Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench,
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads—where latency, cost, and bottlenecks arise—remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference.
Recent research identifies six key properties distinguishing agentic workloads from traditional LLM serving:
To address these challenges, design explorations such as task-aware serving, communication-aware placement, state offloading, and tool-result caching have demonstrated significant improvements, including 29–40% latency reduction, 4.5x speedups in placement, and 35.2% reduction in redundant search calls.
The paper studies the systems-level behavior of LLM-powered agentic applications, where an LLM does not merely answer a one-shot prompt but drives a long-running workflow involving planning, tool calls, external APIs, code execution, and persistent state. This changes the serving problem from stateless, request-isolated inference to session- or task-oriented workloads with heterogeneous components, variable durations, and interdependencies. Existing inference-serving assumptions—short-lived requests, independent batches, predictable time-to-token behavior, and GPU-centric bottlenecks—may therefore be insufficient for understanding where end-to-end latency, cost, and throughput limits actually arise.
AgentSysBench is introduced as a benchmark and characterization framework for these agentic workloads. Rather than evaluating only model accuracy or single-request inference metrics, it profiles the full execution path of agentic tasks and separates the contributions of LLM inference, tool and environment latency, state management, synchronization, queuing, and resource contention. The resulting characterization is meant to expose which stages dominate wall-clock time and cost under different workload patterns, and how agentic execution differs from conventional inference in terms of burstiness, tail latency, resource utilization, and scalability.
The work matters because agentic systems are becoming a primary deployment target for LLMs, but production serving stacks are still largely optimized for request/response inference. By providing empirical evidence about workload structure and bottlenecks, the paper gives system designers a basis for more appropriate scheduling, autoscaling, caching, session affinity, preemption, and cost-control mechanisms. In short, it helps shift serving-system design from assumptions inherited from chat-style inference toward workload-aware architectures for long-running, tool-mediated LLM agents.