In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $n^* \approx 156$--$168$ tokens, HBM weight streaming dominates---cost attaches to $activated replicas$, not tokens
TEMPO is a makespan-aware dispatcher for expert-parallel Mixture-of-Experts (MoE) serving that optimizes load balancing by modeling actual expert execution time rather than relying on proxy metrics like token or expert counts. It solves the per-batch dispatch as a fixed-charge makespan problem in milliseconds off the critical path, integrating with SGLang via a fused in-graph kernel for zero-overhead dispatch.
Unlike existing dispatchers such as EPLB, LPLB, or METRO, which assume linear relationships between time and either token count or activated experts, TEMPO accounts for the two-regime structure of expert costs: memory-bound (where loading weights dominates) and compute-bound (where GEMM dominates). This allows TEMPO to balance across mixed regimes within a single batch, staying within 1% of the best fixed baseline and winning by up to 15.5% where regimes mix.
Key advantages include: Regime-Awareness: It dynamically adapts to both memory-bound and compute-bound phases, whereas fixed proxies fail in one or the other. Performance Gains: TEMPO delivers +38% to +70% request throughput on decode workloads and reduces median inter-token latency by roughly half compared to adaptive dispatch methods. * Efficiency: The solver runs out-of-process with minimal overhead, avoiding the collective communication costs embedded in in-graph solvers like LPLB.
TEMPO examines a subtle but important failure mode in expert-parallel (EP) serving of mixture-of-experts models: because each layer must synchronize across GPUs, the effective latency of a request is governed by the slowest device, making load balancing a makespan-minimization problem rather than merely an average-load problem. Existing EP load-balancing approaches typically assume that expert execution time scales linearly with a single proxy metric—either the number of tokens dispatched to an expert, as in EPLB/LPLB/UltraEP, or the number of activated experts, as in METRO. The paper argues that this assumption is too coarse: it can leave individual GPUs under- or over-loaded even when the global token or expert counts appear balanced.
The central empirical contribution is a regime-aware characterization of expert execution cost on two datacenter GPU generations. The authors show that expert time is not linearly determined by token count or activated-expert count alone. Below a critical batch size of roughly \(n^* \approx 156\)–\(168\) tokens, expert computation is memory-bound: HBM weight streaming dominates, so the cost is driven primarily by how many activated expert replicas must be fetched, not by how many tokens are processed. Above that threshold, the system shifts toward a compute-bound regime where token-level work becomes more important. TEMPO uses this distinction to balance EP workloads in a makespan-aware way, accounting for both memory-bound replica activation and compute-bound token processing rather than relying on a single static balancing metric.
This matters because many practical MoE serving workloads—especially low-batch, latency-sensitive, or interactive inference—operate in or near the memory-bound regime where conventional token-count balancing can be misleading. By exposing the transition between memory- and compute-bound behavior and incorporating it into the load-balancing objective, TEMPO offers a more hardware-faithful model for EP scheduling. The result is a design principle for MoE serving systems: effective expert-parallel load balancing should be regime-aware and makespan-oriented, balancing not just “how much work” is assigned, but what that work actually costs on the target GPU.