Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench,

Topological visualization of From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
Brave API

Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads—where latency, cost, and bottlenecks arise—remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference.

Recent research identifies six key properties distinguishing agentic workloads from traditional LLM serving:

  • Heavyweight and Stateful Execution: Agentic workloads are long-running and accumulate significant live session state, with non-LLM components (sandboxes, retrieval engines) dominating latency in many applications.
  • Cross-Stack Heterogeneity: Workloads combine components with divergent resource profiles (GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes), causing up to 32x latency divergence and head-of-line blocking.
  • Shifting Bottlenecks: The dominant bottleneck is not static but shifts dynamically across requests, models, and deployments based on workload and system factors.
  • Long Idle Intervals: Production sessions often hold state idle for minutes to hours between active steps, complicating resource management.
  • Control-Plane Tax: Auxiliary LLM calls and context overhead from tool schemas and observations consume significant compute and context capacity.
  • Cross-Request Redundancy: Production traces reveal heavy redundancy in search queries and web fetches, exposing large opportunities for caching.

To address these challenges, design explorations such as task-aware serving, communication-aware placement, state offloading, and tool-result caching have demonstrated significant improvements, including 29–40% latency reduction, 4.5x speedups in placement, and 35.2% reduction in redundant search calls.

Generated Sep 2, 2026
Open-Weights Reasoning

The paper studies the systems-level behavior of LLM-powered agentic applications, where an LLM does not merely answer a one-shot prompt but drives a long-running workflow involving planning, tool calls, external APIs, code execution, and persistent state. This changes the serving problem from stateless, request-isolated inference to session- or task-oriented workloads with heterogeneous components, variable durations, and interdependencies. Existing inference-serving assumptions—short-lived requests, independent batches, predictable time-to-token behavior, and GPU-centric bottlenecks—may therefore be insufficient for understanding where end-to-end latency, cost, and throughput limits actually arise.

AgentSysBench is introduced as a benchmark and characterization framework for these agentic workloads. Rather than evaluating only model accuracy or single-request inference metrics, it profiles the full execution path of agentic tasks and separates the contributions of LLM inference, tool and environment latency, state management, synchronization, queuing, and resource contention. The resulting characterization is meant to expose which stages dominate wall-clock time and cost under different workload patterns, and how agentic execution differs from conventional inference in terms of burstiness, tail latency, resource utilization, and scalability.

The work matters because agentic systems are becoming a primary deployment target for LLMs, but production serving stacks are still largely optimized for request/response inference. By providing empirical evidence about workload structure and bottlenecks, the paper gives system designers a basis for more appropriate scheduling, autoscaling, caching, session affinity, preemption, and cost-control mechanisms. In short, it helps shift serving-system design from assumptions inherited from chat-style inference toward workload-aware architectures for long-running, tool-mediated LLM agents.

Generated Sep 2, 2026
Sources