arXiv:2609.33847v1 Announce Type: new Abstract: Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch schedulers such as Slurm. Growing demand for shared computational resources increases the need for interoperability between these environments. This study u
The paper "Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is Required" (arXiv:2609.33847v1, published September 27, 2026) by Mindaugas Macernis argues that pre-execution benchmarking is essential when deploying scientific applications in containers on HPC systems.
The study compares Podman-based containerized execution against on-system execution for quantum chemistry software, highlighting that performance depends on software compilation, startup overhead, and execution configuration. Key findings include:
The paper addresses the growing overlap between cloud and HPC workloads, where the same application may need to run on Kubernetes-managed cloud resources and Slurm-managed HPC clusters. Although containerization and portable runtimes make it easier to move workloads across these environments, the authors emphasize that portability is not equivalent to performance equivalence. Heterogeneous CPUs, accelerators, interconnects, memory hierarchies, file systems, and scheduling semantics can cause a portable workload to run at widely different efficiency depending on where it is placed.
Its central contribution is a performance-versus-portability framing for heterogeneous HPC environments. The study argues that static assumptions—such as choosing a node based on nominal hardware specs, cluster reputation, or scheduler policy—are insufficient when workloads are moved across cloud and HPC stacks. Instead, it advocates pre-execution benchmarking as a required step before job submission or placement. By measuring or estimating workload behavior on candidate resources ahead of full execution, operators can identify suitable nodes, avoid pathological placements, and make scheduling decisions that balance portability with actual throughput, latency, or resource utilization.
This matters because it provides a practical mechanism for interoperability between cloud orchestration and traditional HPC scheduling. As shared computational resources become more common, platforms need a way to decide where portable workloads should run without sacrificing performance. The paper’s insight is that benchmarking should be treated as part of the execution pipeline rather than a post-hoc diagnostic: it can guide resource selection, improve job completion times, reduce wasted allocations, and help developers understand the performance cost of maintaining portability across heterogeneous environments.