Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like description of an AI computation and automatically distributes it across a multi-GPU server, with no device assignments, sharding annotations, or communication operations written by the programmer. Einsummable
Einsummable is a prototype system that automatically distributes AI computations across multi-GPU servers by treating every operation as a relational join followed by an aggregation over tensor relations. It accepts a PyTorch-like description of an AI computation and automatically generates the necessary device assignments, sharding, and communication without programmer intervention.
The system uses join-agg specs to expose possible decompositions for each operation and employs an optimizer to select plans that minimize a communication-cost proxy. This approach allows Einsummable to discover decomposition plans that mesh-based auto-parallelizers cannot, synthesizing topology-aware exchange programs rather than relying on canned collective libraries like NCCL.
Experimental results on an eight-GPU A100 server demonstrate that Einsummable achieves a geometric-mean runtime of 8.97 ms for LLaMA transformer blocks, outperforming hand-tuned PyTorch (13.65 ms) and vLLM (14.87 ms).
Einsummable is a prototype automatic multi-GPU parallelization system for AI workloads. It accepts a PyTorch-like description of a computation and, without programmer-provided device assignments, sharding annotations, or explicit communication primitives, derives a distributed execution plan across the GPUs in a server. The core idea, suggested by the title, is to treat AI kernels through a “join” lens: a kernel is viewed as an index-space operation that combines, reduces, or contracts tensor fragments according to the semantics of the computation. This framing lets the system reason about where data must reside, which GPU fragments are needed for each operation, and what communication is required to satisfy those dependencies.
The paper’s main contribution is a unified abstraction for automatic device placement, tensor sharding, and collective communication. Rather than forcing the programmer to choose between data parallelism, tensor parallelism, pipeline parallelism, or hybrid strategies, Einsummable infers the physical execution plan from the logical tensor program. By expressing computations in a PyTorch-like, likely einsum-compatible form, the system can exploit algebraic structure—such as contraction order, reduction locality, and fragment reuse—to reduce redundant data movement and avoid unnecessary all-gathers, all-reduces, or point-to-point transfers. The result is a more declarative model of multi-GPU execution, where the programmer specifies what computation to perform and the system determines how to partition and schedule it.
This matters because manual sharding and device placement are among the most brittle and error-prone parts of large-scale AI systems. Hand-written parallelism strategies are tightly coupled to model structure, hardware topology, and memory constraints, making them difficult to port and maintain as models and interconnects evolve. Einsummable’s approach could lower the barrier to efficient multi-GPU training and inference by making distribution a compiler/runtime concern rather than a programmer burden. As a prototype, it is best read as a demonstration of an automatic parallelism design space rather than a production-ready framework, but its significance lies in showing that high-level tensor semantics can be used to derive much of the low-level communication and placement logic required for modern multi-GPU AI workloads.
The material presents Einsummable, a prototype system for automatically parallelizing AI computations across the GPUs in a multi-GPU server. It accepts a high-level, PyTorch-like description of a computation and distributes it without requiring the programmer to write device assignments, sharding annotations, or explicit communication operations. The core idea is to express AI kernels in an einsum-like or tensor-contraction form, treating them as join-like operations over indexed tensors so that the system can reason about their data dependencies, intermediate tensor shapes, and parallelization opportunities in a uniform way.
A key contribution is the use of this algebraic/tensor-contraction abstraction as a basis for automatic distributed execution planning. Instead of relying on manually annotated sharding strategies or per-kernel communication code, Einsummable can infer how to partition computation and data across GPUs, assign work to devices, and insert the necessary communication. The “every kernel is a join” framing suggests that a broad class of AI operations can be handled by the same parallelization machinery, allowing the compiler to consider the whole computation rather than treating each kernel as an isolated, manually parallelized unit.
This matters because multi-GPU distribution remains a major systems bottleneck in AI development: manual sharding and collective communication are complex, error-prone, and tightly coupled to hardware topology. An automatic approach like Einsummable could reduce the engineering cost of scaling AI workloads, improve portability across different multi-GPU configurations, and make it easier to express and execute large computations without low-level distributed-programming expertise. As a prototype, its main value is in demonstrating a path toward compiler-driven, annotation-free multi-GPU parallelism for AI systems.