arXiv:2609.24802v1 Announce Type: new Abstract: Graph message passing offers a common way to express learning algorithms, physical simulations, and numerical solvers. Efficient execution depends on interaction structure and data movement, which can be obscured when a program is expressed as a sequence of tensor operations. On memory-constrained systems such as laptops, materializing connectivity

Topological visualization of Tiga: Compiling Graph Message Passing at Scale
Brave API

Tiga (Target-Independent Graph Acceleration) is a just-in-time compiler for graph message-passing programs that separates program semantics from hardware-specific execution, enabling efficient processing on CPU and NVIDIA GPU backends. It addresses memory constraints by streaming graph partitions from disk through host memory with paged device buffers, allowing execution on graphs with over one billion edges on memory-limited hardware.

Key capabilities include generated relation fusion to reduce intermediate storage, distributed execution via halo exchange and rank ownership, and integration with PyTorch autograd for differentiable workloads. While it demonstrates significant reductions in runtime and device memory usage for generated relations, distributed scaling on heterogeneous devices currently faces communication overhead challenges that can negate speedups compared to single-GPU execution.

Generated 12d ago
Open-Weights Reasoning

Tiga: Compiling Graph Message Passing at Scale presents a compiler-oriented approach to executing graph message passing, a primitive underlying graph neural networks, physical simulations, and numerical solvers. The paper argues that when such programs are written as ordinary sequences of tensor operations, their essential graph structure—neighborhoods, interaction patterns, and irregular data dependencies—can become hidden. That obscuration makes it harder for compilers and runtime systems to reason about the dominant costs in practice: memory traffic, cache behavior, and the movement of values across graph elements.

The key contribution is a compilation strategy, named Tiga, that treats graph message passing as a first-class computational pattern rather than a stack of opaque tensor kernels. By exposing the interaction structure of the computation, Tiga can optimize data movement more directly and avoid unnecessary materialization of large connectivity structures. This is especially important on memory-constrained platforms such as laptops, where explicitly constructing or storing full graph adjacency representations can dominate memory usage and limit problem size.

The work matters because it addresses a practical gap between high-level graph algorithms and efficient low-level execution. If graph message passing can be compiled with awareness of both graph topology and tensor-level data movement, systems can support larger and more expressive workloads without requiring high-memory accelerators. More broadly, the paper points toward a useful middle ground between graph-specific frameworks and general tensor compilers, with implications for scalable learning systems, scientific computing, and edge deployment.

Generated 12d ago
Sources