arXiv:2609.24639v1 Announce Type: new Abstract: Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill
The paper Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference (arXiv:2609.24639v1, submitted September 21, 2026) by Mingyuan Yan et al. introduces an analytical framework to optimize the tradeoff between serving capacity and power consumption in disaggregated AI inference.
Problem and scope. The paper addresses a growing operational bottleneck in large-scale AI inference: power availability. As inference fleets scale, datacenter and edge deployments may be limited not only by GPU count or memory capacity, but also by electrical budgets, cooling constraints, and grid or facility power caps. The work focuses on prefill–decode (PD) disaggregated serving, an increasingly common architecture that separates the compute-intensive prompt-processing phase from the memory- and bandwidth-sensitive token-generation phase onto distinct resource pools. This disaggregation improves serving efficiency, but it complicates provisioning because operators must decide how many prefill nodes, decode nodes, or accelerator instances to allocate while simultaneously satisfying throughput, latency, and power constraints.
Key contribution and insights. The central contribution is an analytical framework for power-aware provisioning of PD-disaggregated inference systems. Rather than treating power as a secondary cost metric, the paper models power consumption as a first-class constraint in the capacity-planning problem. It connects workload characteristics—such as request mix, sequence lengths, batch behavior, and service-level objectives—to the required numbers of prefill and decode resources and to the resulting power draw. The resulting analysis helps identify the tradeoff between prefill capacity and decode capacity under a fixed power budget: too much prefill capacity may leave decode underprovisioned for long-running generation, while too much decode capacity may waste power on memory-resident but underutilized resources. The work therefore provides a principled way to choose the PD resource split, or to evaluate how optimal provisioning shifts as workload patterns, SLOs, and power caps change.
Why it matters. This matters because modern inference systems are increasingly sized under real power limits, not just idealized compute availability. A power-blind provisioning approach can overestimate feasible throughput, underestimate energy cost, or cause capacity plans that cannot be deployed in practice. By providing an analytical treatment of PD disaggregation under power constraints, the paper offers operators and system designers a clearer basis for capacity planning, cost optimization, and sustainable scaling. It also highlights a broader architectural insight: disaggregated inference gains are not only about improving utilization or latency isolation, but also about making power a manageable, optimizable dimension of the serving stack.