arXiv:2609.25949v1 Announce Type: cross Abstract: Optical Circuit Switching (OCS) offers high bandwidth density and energy efficiency for LLM training, but incurs a non-negligible reconfiguration delay. Prior work typically schedules optical circuit switches independently of compute, using aggregate traffic demand to determine which circuits to provision and when. We argue that this separation cr
Flux (arXiv:2609.25949) is a workload-aware circuit scheduling framework for Large Language Model (LLM) training that optimally schedules Optical Circuit Switches (OCS) by integrating compute and communication timelines. Unlike traditional schedulers that use aggregate traffic matrices and ignore the temporal structure of workloads, Flux formulates scheduling as a Mixed Integer Linear Program (MILP) to maximize circuit reuse and amortize reconfiguration delays behind active compute phases.
This approach significantly enhances efficiency by: Reducing training iteration time by up to 10× compared to traditional periodic schedulers. Lowering peak Network Interface Card (NIC) buffer requirements by more than three orders of magnitude. * Overlapping communication and reconfiguration overhead with computation, thereby eliminating stalls caused by mismatched circuit availability.
Flux: Optimal Scheduling of Optical Circuit Switches for LLM Training examines how optical circuit switching (OCS) can be better integrated into large-scale LLM training clusters. OCS can provide very high bandwidth at lower energy cost than many electronic switching fabrics, but it has a non-negligible reconfiguration delay when changing optical paths. The paper’s central critique is that prior approaches often treat the network as an independent provisioning problem: they estimate aggregate communication demand and decide which circuits to set up without tightly coupling that decision to the actual compute schedule of the training job. This separation can be suboptimal because LLM training workloads have structured, phase-dependent communication patterns—especially around collectives such as all-reduce—and the timing of network reconfiguration directly affects whether compute units remain busy.
The key contribution of Flux is to co-schedule OCS reconfiguration with the training workload rather than treating them as separate layers. It frames the problem as an optimal or near-optimal scheduling problem in which the system decides not only which optical circuits to provision, but when to switch them, taking into account reconfiguration latency, communication phases, and the opportunity to overlap network setup with useful compute work. The main insight is that OCS performance is not determined solely by raw bandwidth or average traffic volume; it depends on how well the optical topology is aligned in time with the workload’s communication schedule. By making the scheduler workload-aware and compute-aware, Flux can reduce idle time, avoid misaligned circuit provisioning, and better exploit the high-bandwidth, energy-efficient properties of OCS.
This work matters because it addresses a practical bottleneck in scaling LLM training with emerging optical interconnects. As clusters grow, communication overhead and network energy become first-order concerns, and OCS is an attractive but timing-sensitive option. The paper argues that simply adding optical circuits to an existing training cluster is not enough: the scheduler must be redesigned to account for reconfiguration dynamics. For systems researchers, this bridges datacenter networking, optical switching, and AI workload scheduling, suggesting that future high-performance LLM training infrastructure should treat network topology control as part of the execution schedule rather than as a static or loosely coupled provisioning decision.