arXiv:2610.01380v1 Announce Type: new Abstract: GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparison

Topological visualization of GPU-Initiated Communication: Dissecting Down to the Bone
Brave API

GPU-Initiated Communication: Dissecting Down to the Bone (arXiv:2610.01380) is a October 2026 paper by Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay, and Didem Unat that analyzes the performance characteristics and optimizations of GPU-initiated communication at the GPU-NIC boundary.

  • Core Mechanism: The study dissects how GPU threads post RDMA operations directly to the NIC, a pattern underlying libraries like NVSHMEM, NCCL GIN, and DeepEP for fine-grained, latency-critical communication in Mixture-of-Experts (MoE) models.
  • Methodology: The authors introduce mini-gda and mini-proxy, minimal transport implementations that isolate hardware mechanism costs from library overhead, measuring them across NVIDIA H100, H200, B200, and GB200 platforms.
  • Key Findings: Achieving the 260 M msg/s ceiling on InfiniBand requires doorbell batching and queue parallelism, which incur resource costs such as reduced GPU block residency. The paper reveals that library design choices significantly impact latency and message rate, with all-to-all traffic losing 59% of its NIC message rate at approximately 3,000 active connections.
  • Contributions: The work provides a mechanism-level description of the GPU-NIC boundary and releases mini-gda, mini-proxy, and a microbenchmark suite as open source to serve as a performance oracle for communication library developers.
Generated 1d ago
Open-Weights Reasoning

The paper provides a low-level analysis of GPU-initiated communication, the mechanism by which GPU threads directly issue RDMA operations to the NIC instead of relying on CPU-side orchestration. It focuses on the libraries that make this capability practical—NVSHMEM, NCCL GIN, and DeepEP—and on the workloads that depend on it, especially the fine-grained, latency-sensitive messages exchanged by Mixture-of-Experts models during all-to-all and expert-parallel communication. The central concern is that these systems are increasingly important but remain poorly documented: much of their behavior is visible only in source code, benchmarks, or implementation details rather than in a clear performance model.

Its main contribution is a “down to the bone” dissection of how GPU-initiated communication actually works and where its performance comes from. Rather than treating NVSHMEM, NCCL GIN, and DeepEP as black boxes, the paper compares their mechanisms and optimizations, exposing how GPU threads post RDMA requests and how queue management, doorbell signaling, memory registration, completion handling, and synchronization affect latency and throughput. It highlights the trade-offs among API abstraction, batching, message size, and hardware interaction, and identifies which implementation choices matter most for small-message latency versus larger transfers.

This matters because MoE systems are increasingly bottlenecked by communication latency rather than raw bandwidth, and GPU-initiated paths can remove CPU round-trips from the critical path. By documenting performance characteristics and comparing libraries, the work gives systems researchers and practitioners a clearer basis for choosing, tuning, or extending GPU communication stacks. It also helps explain why seemingly similar libraries behave differently under fine-grained MoE traffic and points to design principles for future low-latency interconnects, runtimes, and collective communications.

Generated 1d ago
Sources