arXiv:2609.29707v1 Announce Type: new Abstract: LLM agents that iteratively reason, plan, and invoke tools create workload profiles fundamentally different from single-pass inference, yet how their energy consumption is distributed across hardware components and workload phases remains poorly understood. Characterizing these workloads therefore requires simultaneous visibility into both component

Topological visualization of Where Does the Energy Go? Profiling LLM Agent Inference on Blackwell GPUs
Brave API

LLM agent inference on Blackwell GPUs consumes significantly more energy than single-pass serving due to lack of batching, context growth, and tool-induced idle periods. A full-stack study using Qwen3.8-27B on NVIDIA RTX PRO 6000 Blackwell GPUs revealed that sequential agent workloads consume 63× more system energy per output token than saturated continuous serving.

Key findings include: GPU telemetry is insufficient, missing 41–45% of total system energy; non-GPU components (CPU, DRAM, fans) account for the remainder. Extended thinking modes generate 21–75% more tokens, increasing total energy by volume rather than per-token inefficiency (per-token energy differs by <1%). Continuous batching improves energy efficiency by 3.2× (from 1 to 16 requests per second) as GPU power plateaus while throughput increases. Context accumulation causes per-turn energy to increase over time due to longer prefill durations, suggesting context management techniques like summarization are critical for optimization.

Generated 7d ago
Open-Weights Reasoning

This paper examines the energy behavior of LLM agent inference on Blackwell GPUs, focusing on the fact that agentic workloads are structurally different from conventional single-pass LLM serving. Rather than a one-shot prompt-to-response flow, agents repeatedly reason, plan, call tools, update state, and return to the model, producing multi-phase, bursty, and often non-stationary execution patterns. The central contribution is a profiling approach that provides simultaneous visibility into where energy is consumed—across hardware components such as the GPU, memory, CPU, and interconnect—and how that consumption is distributed across workload phases such as prompt processing, decoding, tool execution, synchronization, and idle or waiting periods.

The key insight is that agent inference energy cannot be reduced to simple proxies like token count, model size, or total wall-clock time. The same model may consume energy very differently depending on how much time is spent in compute-heavy reasoning phases versus I/O-bound tool calls, how long context grows across iterations, how often the GPU is underutilized during external calls, and how much overhead is introduced by orchestration and synchronization. By profiling on Blackwell GPUs, the work highlights that energy attribution requires phase-aware and component-level measurement, because different parts of the agent loop can stress different parts of the system and dominate the cost profile in different ways.

This matters because agentic LLM workloads are moving from research prototypes toward production systems where energy, cost, latency, and sustainability are tightly coupled. Fine-grained profiling helps operators and system designers identify where efficiency is lost—whether in GPU utilization, memory traffic, tool orchestration, context management, or hardware provisioning—and provides a basis for better scheduling, batching, caching, and serving strategies. More broadly, it gives the systems community a measurement foundation for evaluating not just model performance, but the operational energy profile of increasingly autonomous LLM-driven applications.

Generated 7d ago
Sources