Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the prof

Topological visualization of Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
Brave API

Phase-Decoupled, Model-Calibrated Power Control is a system for disaggregated LLM serving that uses distinct power control mechanisms for the compute-bound prefill phase and the memory-bound decode phase, calibrated specifically for each model and engine stack rather than relying on generic GPU-class profiles.

The system improves upon NVIDIA’s Max-Q inference profile by applying an SM-clock window to prefill GPUs (bounded by latency guarantees) and a calibrated power cap to decode GPUs (placed just above the throughput cliff, allowing the GPU’s internal power manager to handle frequency scaling). This phase-decoupled approach yields a Pareto improvement over Max-Q, delivering +20.4% tokens/J at only +3.5% mean end-to-end latency on an 8x B200 node serving Qwen3-Coder-480B, compared to Max-Q’s +8.6% tokens/J at +5.2% latency.

Key features include: Per-Combination Calibration: Optimal power settings are measured for the specific (model, quantization, engine, hardware) combination, not just the GPU class. Operating-Mode Ladder: Provides selectable modes (PERF, BAL, EFF) that trade latency for energy efficiency while strictly guarding tail latency SLOs. * Elimination of Reactive Overshoot: By capping the flat, memory-bound decode power continuously, the system avoids the reactive overshoot issues that led prior work (like POLCA) to reject power capping for LLM inference.

Generated 22d ago
Open-Weights Reasoning

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving examines how GPU power management should evolve once LLM inference is deployed in prefill/decode-disaggregated datacenters, where available electrical power—not peak compute or memory capacity—often becomes the binding constraint on serving scale. The paper begins with a production-relevant empirical finding: applying NVIDIA’s Max-Q inference profile to a disaggregated B200 system yields only a modest realized energy-efficiency gain of about +8.6% tokens/J, and that gain is highly model-dependent. More importantly, the profile introduces a mean end-to-end latency penalty of roughly +5.2%, a cost that throughput-only or phase-isolated evaluations can easily hide. This motivates a more nuanced view of power control: a single vendor-provided profile is not sufficient for modern disaggregated serving stacks.

The key contribution is a power-control method that decouples the prefill and decode phases and calibrates power budgets to the specific model and serving configuration being run. Prefill and decode have different performance bottlenecks: prefill is typically compute-intensive and bursty, while decode is more sensitive to memory bandwidth, batching dynamics, and sustained power draw. A global power cap or one-size-fits-all efficiency profile can therefore shift the bottleneck in ways that degrade tail or mean latency even when aggregate throughput appears acceptable. The proposed approach treats power allocation as a model-aware, phase-specific control problem, aiming to preserve user-facing SLOs while still capturing meaningful energy savings.

This work matters because AI datacenters are increasingly power-limited, and PD disaggregation is becoming a standard architecture for high-throughput, low-latency LLM serving. The paper shows that energy-aware deployment cannot be reduced to enabling a power-saving mode and reporting tokens/J in isolation; it must account for phase asymmetry, model-specific sensitivity, and end-to-end latency behavior. For operators, the practical takeaway is that modest but real efficiency gains are possible, but only if power control is designed around the actual serving pipeline rather than around raw hardware throughput metrics.

Generated 22d ago
Sources