Modern high-performance GPU computations partition tensors into tiles to exploit data reuse and parallelism. Individual tile computations complete earlier than the full tensor computation, creating opportunities to overlap computation and communication. However, a mismatch between computation and communication progress can limit these opportunities. Communication stalls when no data is ready, and

Topological visualization of Entwine: Coordinating Tiled Computation and Fine-Grained Communication across GPUs
Brave API

The provided search context does not contain information regarding a system or paper titled "Entwine."

The context does, however, detail several other approaches to fine-grained computation-communication overlap for tiled GPU workloads:

  • Tile-level Signaling for MoE: A recent design (July 2026) uses persistent producer-consumer kernels coordinated by device-resident signals to overlap expert GEMM computation with all-to-all communication in Mixture-of-Experts models, achieving up to 2.64x speedup on 4-A100 platforms.
  • GPU-Initiated Fused Kernels: An approach (2023) fuses embedding and All-to-All communication into a single persistent GPU kernel, allowing workgroups to issue non-blocking network transactions immediately upon completion, reducing execution time by 31% on average.
  • COMET: A system (2024) that decomposes shared tensors and reschedules GroupGEMM tiles to overlap communication and computation in MoE layers, delivering up to 1.96x speedup on H800/L20 clusters.
  • Iris: A Triton-based framework (July 2026) that provides first-class multi-GPU programming with value- and pointer-based APIs to enable tile-granular data movement directly from registers to remote memory, overcoming the bulk-synchronous limitations of traditional collective libraries.
Generated 23d ago
Open-Weights Reasoning

Entwine addresses a core inefficiency in tiled, multi-GPU tensor computations: computation and communication often advance at different granularities and speeds. In such workloads, tensors are divided into tiles so that GPUs can exploit locality, reuse, and parallelism, and individual tiles may finish well before the full tensor is complete. This creates natural windows in which communication could begin early, but conventional coarse-grained synchronization or communication patterns may miss the opportunity, causing communication to wait for data that is not yet ready or computation to wait for data that has not arrived.

The key insight of the material is that tile-level progress should be treated as a first-class signal for scheduling communication. Entwine proposes coordinating tiled computation with fine-grained inter-GPU communication so that data movement can be issued, reordered, or consumed according to the actual readiness of tiles rather than fixed tensor-level barriers or large collective boundaries. In doing so, it aims to reduce idle time on both compute and communication paths by more tightly coupling when a tile becomes available with when it is transferred or consumed.

This matters because communication overhead is a major limiter of scalability in modern GPU systems, especially for distributed training, tensor-parallel workloads, and other latency-sensitive multi-GPU applications. By exposing and exploiting the temporal irregularity of tile completion, Entwine offers a design direction for runtimes that can improve GPU utilization, lower end-to-end latency, and make better use of high-speed interconnects without requiring programmers to manually manage complex pipelining and synchronization.

Generated 23d ago
Sources