arXiv:2610.01380v1 Announce Type: new Abstract: GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparison
GPU-Initiated Communication: Dissecting Down to the Bone (arXiv:2610.01380) is a October 2026 paper by Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay, and Didem Unat that analyzes the performance characteristics and optimizations of GPU-initiated communication at the GPU-NIC boundary.
The paper provides a low-level analysis of GPU-initiated communication, the mechanism by which GPU threads directly issue RDMA operations to the NIC instead of relying on CPU-side orchestration. It focuses on the libraries that make this capability practical—NVSHMEM, NCCL GIN, and DeepEP—and on the workloads that depend on it, especially the fine-grained, latency-sensitive messages exchanged by Mixture-of-Experts models during all-to-all and expert-parallel communication. The central concern is that these systems are increasingly important but remain poorly documented: much of their behavior is visible only in source code, benchmarks, or implementation details rather than in a clear performance model.
Its main contribution is a “down to the bone” dissection of how GPU-initiated communication actually works and where its performance comes from. Rather than treating NVSHMEM, NCCL GIN, and DeepEP as black boxes, the paper compares their mechanisms and optimizations, exposing how GPU threads post RDMA requests and how queue management, doorbell signaling, memory registration, completion handling, and synchronization affect latency and throughput. It highlights the trade-offs among API abstraction, batching, message size, and hardware interaction, and identifies which implementation choices matter most for small-message latency versus larger transfers.
This matters because MoE systems are increasingly bottlenecked by communication latency rather than raw bandwidth, and GPU-initiated paths can remove CPU round-trips from the critical path. By documenting performance characteristics and comparing libraries, the work gives systems researchers and practitioners a clearer basis for choosing, tuning, or extending GPU communication stacks. It also helps explain why seemingly similar libraries behave differently under fine-grained MoE traffic and points to design principles for future low-latency interconnects, runtimes, and collective communications.