arXiv:2609.33762v1 Announce Type: new Abstract: LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it s

Topological visualization of EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
Brave API

EfficientAgent resolves the inconsistent performance of KV cache offloading in concurrent LLM agents by sizing the host memory tier to match the reuse working set—the total context of all active agents processed between an agent’s turns.

  • Core Problem: Standard offloading often fails because cached state is evicted by other agents' requests before the original agent returns, leading to unnecessary recomputation despite available bandwidth.
  • Solution: The system uses a stack-distance model to estimate the working set from agent histories, ensuring the host tier is large enough to retain reusable prefixes.
  • Mechanism: It employs working-set-aware admission scheduling, which applies load control to host writes: stopping large refills when the tier is undersized to prevent thrashing, and writing everything when it is sufficient.
  • Results: On SWE-bench Verified coding agents, this approach cuts recomputed prompt tokens by 93% and reduces end-to-end time by 39% compared to fixed-tier strategies.
Generated 5d ago
Open-Weights Reasoning

The paper examines why KV-cache offloading, a common memory-management technique in LLM serving, produces inconsistent results under agentic workloads. LLM agents typically resend their full conversation history on every turn, so much of the input is repeated prefix that could be served from cached key-value state rather than recomputed. When GPU memory is exhausted, serving systems often offload less active KV blocks to host memory. The authors focus on the harder regime where many agents run concurrently—especially coding-agent workloads with long contexts, repeated tool calls, and bursty turn structure—where offloading can sometimes relieve memory pressure but can also introduce transfer stalls, eviction churn, and poor cache locality.

The key contribution is an analysis of the conditions under which offloaded KV state is actually useful for concurrent agents. Rather than treating KV offloading as a uniform win, the work identifies how factors such as prefix reuse, request batching, GPU memory availability, host-device bandwidth, and scheduling interact to determine whether offloading improves or degrades performance. From this perspective, the paper frames EfficientAgent as an agent-aware approach to making KV offloading more predictable and effective, by aligning cache placement and serving behavior with the repetitive, stateful nature of multi-turn agent execution.

This matters because agentic LLM serving is moving beyond single-shot inference toward long-lived, memory-intensive applications. If offloading behavior is unpredictable, operators may see erratic tail latency and throughput even on identical workloads. By clarifying what makes KV offloading work—or fail—for concurrent agents, the paper provides a practical foundation for building more predictable, memory-efficient LLM serving stacks for agent-centric workloads.

Generated 5d ago
Sources