Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intra-node) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a single compute node is essential for applications in scientific and engineering

Topological visualization of Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study
Brave API

OpenMP offloading provides a portable, single-codebase solution for intra-node multi-GPU programming across NVIDIA, AMD, and Intel architectures, though performance varies significantly by vendor and implementation strategy.

NVIDIA GPUs: OpenMP offloading achieves comparable performance to vendor-specific libraries (cuBLAS) for BLAS Level 1-2 operations, with close-to-ideal strong scaling in multi-GPU nodes when using task-based implementations. However, for dense matrix-vector multiplication, it may reach only 50-60% of CUDA performance and scales poorly as GPU count increases due to kernel inefficiencies.

AMD GPUs: Performance is generally lower than vendor-specific HIP implementations, achieving only ~10-25% of hipBLAS performance for certain operations. The Infinity Fabric interconnect provides ~1.5x bandwidth over PCIe, aiding data transfer, but compiler support (Cray Clang) lags behind NVIDIA’s NVHPC, particularly for matrix-matrix multiplications.

Intel GPUs (PVC): The Intel GPU Max 1550 shows ~40% higher performance than comparable NVIDIA A100 or AMD MI250X GPUs for specific solvers (e.g., STREAmS compressible solver). OpenMP offloading on Intel hardware demonstrates good intra-node and internode scalability, with strong scaling efficiency around 52% and weak scaling efficiency of 0.88 in multi-node configurations.

Key Optimization Strategies: Data Transfer: Using low-level APIs for CPU-GPU data movement significantly outperforms high-level structured mapping, reducing transfer overhead. Task-Based Scheduling: Distributing work across multiple GPUs via OpenMP taskloop constructs improves performance by up to 57.2% compared to static assignment, enabling better overlap of computation and data transfer. * Compiler Selection: NVHPC is recommended for NVIDIA, while Cray Clang is preferred for AMD; Clang is not recommended for NVIDIA offloading due to poor performance.

Generated Aug 23, 2026
Open-Weights Reasoning

The paper investigates how well OpenMP target offloading supports practical intra-node multi-GPU programming on current heterogeneous supercomputing hardware, using a 3D heat transfer solver as a representative scientific case study. Rather than focusing on a single accelerator ecosystem, it evaluates the OpenMP offload model across NVIDIA, AMD, and Intel GPU architectures, examining whether high-level OpenMP constructs can effectively distribute and execute GPU-resident workloads across multiple devices within one compute node. The heat transfer application provides a realistic but controlled benchmark with substantial parallelism, memory traffic, and inter-iteration dependencies, making it useful for assessing both correctness and performance of multi-GPU offloading.

Its key contribution is a cross-vendor perspective on the practical maturity of OpenMP for multi-GPU workloads, including how portability, compiler support, runtime behavior, and performance scale across different GPU platforms. By applying the same OpenMP-based application to multiple architectures, the study highlights where the standard provides a workable abstraction and where vendor-specific differences still affect code usability, tuning, or efficiency. This is important because many production scientific applications need to run on diverse accelerator systems without maintaining separate code paths for each GPU vendor, and the results help clarify whether OpenMP offloading is becoming a credible general-purpose model for exploiting multiple GPUs within a single node.

Generated Aug 23, 2026
Sources