arXiv:2610.01587v1 Announce Type: new Abstract: Parallel applications are often designed for synchronous, lock-step execution, treating communication stalls as performance hazards. Yet, in a communication-light application without frequent synchronization points that alternates between compute-bound memory-bound execution, an MPI communication stall can act as an unintentional relief on memory-ba
ArXiv:2610.01587v1, titled "Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications," was published on October 1, 2026, by Ayesha Afzal, Krishna Manda, and Georg Hager.
The research challenges the traditional view of MPI communication stalls as performance hazards, demonstrating that in communication-light applications, these stalls can unintentionally relieve memory-bandwidth contention. Key findings include:
The paper examines a counterintuitive performance effect in MPI applications: communication stalls, usually treated as overhead, can sometimes improve the behavior of memory-bound computation. It focuses on communication-light parallel workloads that alternate between compute-bound and memory-bound kernels and do not contain dense synchronization points. In such applications, an MPI communication stall can temporarily pause work on one or more ranks, reducing concurrent memory traffic or giving the memory subsystem time to recover before a subsequent memory-intensive phase. The central insight is that, under certain conditions, communication latency acts as an unintentional throttle that relieves memory-bandwidth pressure rather than simply wasting time.
The broader contribution is a reframing of how MPI performance should be analyzed in mixed compute/memory workloads. Rather than viewing all communication stalls as hazards to be eliminated, the material suggests that the interplay between communication timing and kernel type can be exploited to improve effective memory utilization. This matters because many HPC tuning strategies assume that communication should be minimized or perfectly overlapped with computation; recognizing communication as a potential memory-bandwidth management mechanism can inform better kernel ordering, synchronization placement, and communication-pattern design for MPI codes.