Modern high-performance GPU computations partition tensors into tiles to exploit data reuse and parallelism. Individual tile computations complete earlier than the full tensor computation, creating opportunities to overlap computation and communication. However, a mismatch between computation and communication progress can limit these opportunities. Communication stalls when no data is ready, and
The provided search context does not contain information regarding a system or paper titled "Entwine."
The context does, however, detail several other approaches to fine-grained computation-communication overlap for tiled GPU workloads:
Entwine addresses a core inefficiency in tiled, multi-GPU tensor computations: computation and communication often advance at different granularities and speeds. In such workloads, tensors are divided into tiles so that GPUs can exploit locality, reuse, and parallelism, and individual tiles may finish well before the full tensor is complete. This creates natural windows in which communication could begin early, but conventional coarse-grained synchronization or communication patterns may miss the opportunity, causing communication to wait for data that is not yet ready or computation to wait for data that has not arrived.
The key insight of the material is that tile-level progress should be treated as a first-class signal for scheduling communication. Entwine proposes coordinating tiled computation with fine-grained inter-GPU communication so that data movement can be issued, reordered, or consumed according to the actual readiness of tiles rather than fixed tensor-level barriers or large collective boundaries. In doing so, it aims to reduce idle time on both compute and communication paths by more tightly coupling when a tile becomes available with when it is transferred or consumed.
This matters because communication overhead is a major limiter of scalability in modern GPU systems, especially for distributed training, tensor-parallel workloads, and other latency-sensitive multi-GPU applications. By exposing and exploiting the temporal irregularity of tile completion, Entwine offers a design direction for runtimes that can improve GPU utilization, lower end-to-end latency, and make better use of high-speed interconnects without requiring programmers to manually manage complex pipelining and synchronization.