<div>Dear all,</div><div><br /></div><div>I am running CP2K with the OpenCL backend on Intel Data Center GPU Max 1550 (Ponte Vecchio) nodes on SuperMUC-NG Phase 2, offloading only the DBCSR sparse matrix-matrix multiplication library to the GPU while the rest of the workload remains on the CPU.</div><div>
<p dir="ltr">Benchmark results (below) show that the GPU-offloaded runs are consistently <em>slower</em> than the CPU-only runs. I assume that because only DBCSR is offloaded, the host-device data transfer and synchronization overhead outweighs the compute gain from the GPU </p>
<p dir="ltr">Has this behavior of GPU offload underperforming CPU-only execution when only DBCSR is accelerated been reported before? I'd also appreciate pointers to published OpenCL benchmark results for DBCSR/CP2K on Intel GPUs.</p></div><div><br /></div><div>Below are the results </div><img alt="Screenshot From 2026-08-30 18-02-51.png" width="534px" height="276px" src="cid:cae0105b-425d-4dfd-865c-e3baf08466a5" /><div><br /></div><div>Have a nice day.</div><div><br /></div><div>Thanks and Regards,</div><div>Prasanth.</div>

<p></p>

-- <br />
You received this message because you are subscribed to the Google Groups "cp2k" group.<br />
To unsubscribe from this group and stop receiving emails from it, send an email to <a href="mailto:cp2k+unsubscribe@googlegroups.com">cp2k+unsubscribe@googlegroups.com</a>.<br />
To view this discussion visit <a href="https://groups.google.com/d/msgid/cp2k/c0457faf-6ddb-4dd1-b82b-af7ba5e3c9f4n%40googlegroups.com?utm_medium=email&utm_source=footer">https://groups.google.com/d/msgid/cp2k/c0457faf-6ddb-4dd1-b82b-af7ba5e3c9f4n%40googlegroups.com</a>.<br />