arXiv:2609.24628v1 Announce Type: new Abstract: The High-Luminosity LHC (HL-LHC) will demand order-of-magnitude gains in analysis throughput, and increasingly those gains must come from GPUs that are not made by a single vendor. Leadership-class systems such as El Capitan, Frontier and LUMI are built on AMD accelerators, yet the Scikit-HEP analysis stack---and Awkward Array in particular---has gr
arXiv:2609.24628v1 introduces $rawkward$, a Rust-backed kernel engine that enables AMD GPU (ROCm/HIP) support for Awkward Array, addressing the HL-LHC era's need for vendor-agnostic hardware acceleration beyond NVIDIA's CUDA.
The research identifies that a naive port of CUDA kernels to HIP results in 5–10× performance losses on irregular kernels due to AMD’s 64-lane wavefronts, higher register pressure, and expensive divergence compared to NVIDIA’s 32-thread warps. To close this gap, the authors employ HIP-specific optimization patterns including loop flattening, 128-bit vectorized loads, and splitting fused kernels, which recover CUDA-class performance without altering the public API.
Key performance findings include: GPU Speedups: On a two-socket AMD Instinct MI210 node, the backend achieves speedups from 1.03× (bandwidth-bound sum) to 12.5× (count) over 128 EPYC 7763 CPU cores. CPU Optimization: The Rust CPU kernels match or beat incumbent C++ kernels, with a geometric-mean runtime ratio of 0.37×. * Scalability: The solution ensures that leadership-class systems like El Capitan, Frontier, and LUMI (built on AMD accelerators) can effectively support the Scikit-HEP analysis stack.
This material addresses a practical but consequential gap in the Scikit-HEP ecosystem: high-performance analysis code built on Awkward Array has historically been optimized around NVIDIA CUDA, while next-generation leadership-class HPC systems for the HL-LHC era—such as El Capitan, Frontier, and LUMI—are based on AMD GPUs. The paper frames the need to support AMD accelerators through ROCm/HIP so that the same analysis workflows can run on the hardware that will be available for large-scale LHC data processing. In doing so, it positions AMD support not as a niche portability exercise, but as a requirement for the order-of-magnitude throughput improvements expected at the HL-LHC.
Its central contribution is a technical pathway for bringing Awkward Array’s GPU-accelerated capabilities to AMD platforms. This involves adapting CUDA-based kernels and runtime assumptions to HIP, handling backend-specific memory and launch semantics, preserving the ergonomic Python API, and validating correctness and performance on AMD hardware. The key insight is that vendor-neutral acceleration can be achieved by aligning with the ROCm/HIP model already used across leadership systems, while still maintaining a common Scikit-HEP interface. This reduces the risk that analysis software becomes stranded on a single accelerator vendor as HPC architecture diversifies.
It matters because Awkward Array is a foundational data structure for modern HEP analysis, and its performance characteristics influence how efficiently large datasets can be processed. By enabling AMD GPUs, the work helps the community exploit emerging systems, lowers deployment friction for experiments and collaborations, and supports a more resilient, multi-vendor software stack. For the HL-LHC era, where data volumes and analysis complexity will both grow, such portability is essential to ensuring that computational gains are broadly accessible rather than contingent on proprietary hardware.