Wide-area bandwidth per unit of GPU compute falls every hardware generation: in compute-intensity-ratio terms (CIR, bytes per FLOP), the gap between on-package memory and the conventional WAN is four to five orders of magnitude, widening at roughly 12-19% per year. Position papers - including our own - argued this makes elastic optical wide-area capacity necessary for cross-site AI inference. Revi
Distributed AI inference requires more wide-area bandwidth when agentic workflows involve multi-step reasoning chains that cause KV cache to accumulate continuously without resetting. This accumulation creates a moving bandwidth floor that compounds with every agent interaction, transforming a modest initial cache (e.g., 10 GB) into a significant, persistent data movement requirement.
Consequently, standard infrastructure is insufficient, necessitating a purpose-built "AI WAN" with dedicated coherent optical transport. This architecture provides substantially more total capacity than standard deployments to handle the high-throughput, low-latency connectivity demands of heterogeneous compute nodes across data center campuses or regions.
The paper examines a central tension in cross-site AI inference: as GPU generations improve, the amount of wide-area network bandwidth available per unit of compute keeps falling. It frames the issue using compute-intensity-ratio terms—bytes moved per FLOP—where the gap between on-package memory bandwidth and conventional WAN bandwidth spans four to five orders of magnitude and is widening by roughly 12–19% per year. This trend has led some position work, including the authors’ earlier arguments, to claim that elastic optical wide-area capacity is essential for distributed inference. The paper revisits that claim by asking when distributed AI inference actually requires more WAN bandwidth, and what can be done instead through coordinated design across optical transport, packet networking, and inference software.
Its main contribution is a co-design evaluation that treats WAN bandwidth not as a single fixed constraint, but as a variable shaped by workload, model architecture, transport, and software choices. The analysis distinguishes optical levers, such as elastic optical circuits and bandwidth provisioning, from packet-level mechanisms that improve effective throughput or utilization, and from software-level techniques that reduce cross-site traffic, such as model partitioning, quantization, batching, and cache or state management. The key insight is that the need for additional wide-area capacity is conditional: some distributed-inference regimes are genuinely bandwidth-bound, especially when large model shards, long-context state, or high token rates must be moved across sites, while others can be made viable by tighter co-design rather than by simply adding more WAN bandwidth.
This matters because it refines the common narrative that cross-datacenter AI inference automatically demands ever-larger optical WAN investments. Instead, the paper offers a more nuanced planning framework for AI infrastructure operators: it identifies when elastic optical capacity is the right answer, when packet and software optimizations can defer or reduce the need for it, and when the economics of cross-site inference may still be unfavorable. For datacenter architects, optical network planners, and AI systems researchers, the work provides a useful set of design criteria for aligning model deployment strategies with realistic wide-area bandwidth constraints.