[CP2K-user] [CP2K:22415] DBCSR-only GPU offload (OpenCL) on Intel PVC slower than CPU-only - expected?

ganta.pra...@gmail.com ganta.prasanthbabu at gmail.com
Sun Aug 30 16:21:18 UTC 2026


Dear all,

I am running CP2K with the OpenCL backend on Intel Data Center GPU Max 1550 
(Ponte Vecchio) nodes on SuperMUC-NG Phase 2, offloading only the DBCSR 
sparse matrix-matrix multiplication library to the GPU while the rest of 
the workload remains on the CPU.

Benchmark results (below) show that the GPU-offloaded runs are consistently 
*slower* than the CPU-only runs. I assume that because only DBCSR is 
offloaded, the host-device data transfer and synchronization overhead 
outweighs the compute gain from the GPU 

Has this behavior of GPU offload underperforming CPU-only execution when 
only DBCSR is accelerated been reported before? I'd also appreciate 
pointers to published OpenCL benchmark results for DBCSR/CP2K on Intel GPUs.

Below are the results 
[image: Screenshot From 2026-08-30 18-02-51.png]

Have a nice day.

Thanks and Regards,
Prasanth.

-- 
You received this message because you are subscribed to the Google Groups "cp2k" group.
To unsubscribe from this group and stop receiving emails from it, send an email to cp2k+unsubscribe at googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/cp2k/c0457faf-6ddb-4dd1-b82b-af7ba5e3c9f4n%40googlegroups.com.
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <https://lists.cp2k.org/archives/cp2k-user/attachments/20260830/7cb6df68/attachment-0001.htm>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: Screenshot From 2026-08-30 18-02-51.png
Type: image/png
Size: 416407 bytes
Desc: not available
URL: <https://lists.cp2k.org/archives/cp2k-user/attachments/20260830/7cb6df68/attachment-0001.png>


More information about the CP2K-user mailing list