[CP2K-user] [CP2K:22415] DBCSR-only GPU offload (OpenCL) on Intel PVC slower than CPU-only - expected?
ganta.pra...@gmail.com
ganta.prasanthbabu at gmail.com
Sun Aug 30 16:21:18 UTC 2026
Dear all,
I am running CP2K with the OpenCL backend on Intel Data Center GPU Max 1550
(Ponte Vecchio) nodes on SuperMUC-NG Phase 2, offloading only the DBCSR
sparse matrix-matrix multiplication library to the GPU while the rest of
the workload remains on the CPU.
Benchmark results (below) show that the GPU-offloaded runs are consistently
*slower* than the CPU-only runs. I assume that because only DBCSR is
offloaded, the host-device data transfer and synchronization overhead
outweighs the compute gain from the GPU
Has this behavior of GPU offload underperforming CPU-only execution when
only DBCSR is accelerated been reported before? I'd also appreciate
pointers to published OpenCL benchmark results for DBCSR/CP2K on Intel GPUs.
Below are the results
[image: Screenshot From 2026-08-30 18-02-51.png]
Have a nice day.
Thanks and Regards,
Prasanth.
--
You received this message because you are subscribed to the Google Groups "cp2k" group.
To unsubscribe from this group and stop receiving emails from it, send an email to cp2k+unsubscribe at googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/cp2k/c0457faf-6ddb-4dd1-b82b-af7ba5e3c9f4n%40googlegroups.com.
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <https://lists.cp2k.org/archives/cp2k-user/attachments/20260830/7cb6df68/attachment-0001.htm>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: Screenshot From 2026-08-30 18-02-51.png
Type: image/png
Size: 416407 bytes
Desc: not available
URL: <https://lists.cp2k.org/archives/cp2k-user/attachments/20260830/7cb6df68/attachment-0001.png>
More information about the CP2K-user
mailing list