arXiv:2609.24639v1 Announce Type: new Abstract: Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the appropriate numbers of prefill

Topological visualization of Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference
Brave API

The paper Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference (arXiv:2609.24639v1, submitted September 21, 2026) by Mingyuan Yan et al. introduces an analytical framework to optimize the tradeoff between serving capacity and power consumption in disaggregated AI inference.

  • Core Mechanism: The framework models serving capacity based on the joint distribution of input/output lengths and hardware limits, explicitly capturing the coupling between prefill and decode phases caused by KV-cache reservations and request queueing.
  • Power Modeling: It develops load-dependent power models for prefill and decode instances as functions of normalized serving throughput, allowing for the calculation of average deployment power.
  • Key Output: The combination of these models generates a serving capacity–power Pareto front, enabling service providers to identify minimum-power deployments for required capacities or quantify capacity maintenance under reduced power budgets.
  • Validation: The models were evaluated against production traces from Mooncake and Azure, demonstrating that the analytical approach closely reproduces measured capacity and power, reducing reliance on extensive profiling while supporting grid-friendly inference control.
Generated 12d ago
Open-Weights Reasoning

Problem and scope. The paper addresses a growing operational bottleneck in large-scale AI inference: power availability. As inference fleets scale, datacenter and edge deployments may be limited not only by GPU count or memory capacity, but also by electrical budgets, cooling constraints, and grid or facility power caps. The work focuses on prefill–decode (PD) disaggregated serving, an increasingly common architecture that separates the compute-intensive prompt-processing phase from the memory- and bandwidth-sensitive token-generation phase onto distinct resource pools. This disaggregation improves serving efficiency, but it complicates provisioning because operators must decide how many prefill nodes, decode nodes, or accelerator instances to allocate while simultaneously satisfying throughput, latency, and power constraints.

Key contribution and insights. The central contribution is an analytical framework for power-aware provisioning of PD-disaggregated inference systems. Rather than treating power as a secondary cost metric, the paper models power consumption as a first-class constraint in the capacity-planning problem. It connects workload characteristics—such as request mix, sequence lengths, batch behavior, and service-level objectives—to the required numbers of prefill and decode resources and to the resulting power draw. The resulting analysis helps identify the tradeoff between prefill capacity and decode capacity under a fixed power budget: too much prefill capacity may leave decode underprovisioned for long-running generation, while too much decode capacity may waste power on memory-resident but underutilized resources. The work therefore provides a principled way to choose the PD resource split, or to evaluate how optimal provisioning shifts as workload patterns, SLOs, and power caps change.

Why it matters. This matters because modern inference systems are increasingly sized under real power limits, not just idealized compute availability. A power-blind provisioning approach can overestimate feasible throughput, underestimate energy cost, or cause capacity plans that cannot be deployed in practice. By providing an analytical treatment of PD disaggregation under power constraints, the paper offers operators and system designers a clearer basis for capacity planning, cost optimization, and sustainable scaling. It also highlights a broader architectural insight: disaggregated inference gains are not only about improving utilization or latency isolation, but also about making power a manageable, optimizable dimension of the serving stack.

Generated 12d ago
Sources