arXiv:2610.01784v1 Announce Type: new Abstract: Reducing LLM serving energy does not by itself guarantee lower deployment cost when electricity procurement exposes operators to unfavorable deviations from preset commitments. We study hourly commitments with positive, potentially asymmetric costs for overuse and underuse, and formulate energy-Performance-Aware Commitment Tracking: minimize deviati

Topological visualization of ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving
Brave API

ePACT (arXiv:2610.01784) is an LLM serving system that minimizes the asymmetric deviation costs of preset hourly energy commitments while maintaining request-level Service-Level Objectives (SLOs). Published on October 1, 2026, by You Peng, Youhe Jiang, Chen Wang, and Binhang Yuan, it addresses the fact that reducing physical energy consumption does not guarantee lower deployment costs when electricity procurement involves penalties for both overuse and underuse.

The system employs a two-level controller that adjusts serving capacity and GPU clocks as requests arrive. A global planner updates interval energy targets based on measured consumption and the remaining hourly commitment, while a local decision maker predicts candidate configurations’ energy and completion times to select the optimal setting. This approach allows the system to balance the risk of SLO violations against the costs of deviating from the energy target.

In 24-hour trace simulations for H20 and H200 GPU pools, ePACT reduced asymmetric deviation costs by 73.8% and 75.7% relative to vLLM, while retaining near-vLLM SLO attainment. The system achieved mean absolute hourly deviations of 2.16% and 2.31%, respectively, with 22 and 21 of 24 hours falling within ±5% of their commitments. This demonstrates that adapting serving choices to measured progress toward a commitment can preserve service quality while significantly reducing modeled deviation exposure.

Generated 1d ago
Open-Weights Reasoning

ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving

This paper reframes energy optimization for LLM serving around the economic reality of electricity procurement. It argues that simply minimizing energy use is insufficient when operators must meet preset hourly energy commitments and pay for deviations from them. The work models these commitments as hourly targets with positive, potentially asymmetric penalties for overuse and underuse, and formulates Energy-Performance-Aware Commitment Tracking: the problem of controlling serving behavior so that energy consumption stays close to the committed trajectory while still satisfying performance requirements such as latency or throughput targets.

The key contribution is a system-level perspective that couples energy economics with LLM serving control. Rather than treating energy as a one-sided minimization objective, ePACT treats deviations from the energy commitment as explicit costs and co-optimizes them against application performance. This enables serving mechanisms—such as scheduling, batching, admission control, or other load-shaping decisions—to modulate power draw in a way that avoids both wasteful overconsumption and underutilization that may still incur penalties. The insight is important because LLM workloads are power-intensive and often bursty, making naive energy reduction either economically suboptimal or harmful to user-facing performance.

The material matters because it connects ML systems design with practical data-center cost structures. For operators deploying large language models under power contracts, demand-response arrangements, or grid-related penalties, commitment-aware optimization can reduce total deployment cost without sacrificing service quality. More broadly, it offers a more realistic objective for sustainable AI serving: not merely “use less energy,” but “use the right amount of energy at the right time” under both operational and commercial constraints.

Generated 1d ago
Sources