arXiv:2609.33762v1 Announce Type: new Abstract: LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the same coding-agent workload it s
EfficientAgent resolves the inconsistent performance of KV cache offloading in concurrent LLM agents by sizing the host memory tier to match the reuse working set—the total context of all active agents processed between an agent’s turns.
The paper examines why KV-cache offloading, a common memory-management technique in LLM serving, produces inconsistent results under agentic workloads. LLM agents typically resend their full conversation history on every turn, so much of the input is repeated prefix that could be served from cached key-value state rather than recomputed. When GPU memory is exhausted, serving systems often offload less active KV blocks to host memory. The authors focus on the harder regime where many agents run concurrently—especially coding-agent workloads with long contexts, repeated tool calls, and bursty turn structure—where offloading can sometimes relieve memory pressure but can also introduce transfer stalls, eviction churn, and poor cache locality.
The key contribution is an analysis of the conditions under which offloaded KV state is actually useful for concurrent agents. Rather than treating KV offloading as a uniform win, the work identifies how factors such as prefix reuse, request batching, GPU memory availability, host-device bandwidth, and scheduling interact to determine whether offloading improves or degrades performance. From this perspective, the paper frames EfficientAgent as an agent-aware approach to making KV offloading more predictable and effective, by aligning cache placement and serving behavior with the repetitive, stateful nature of multi-turn agent execution.
This matters because agentic LLM serving is moving beyond single-shot inference toward long-lived, memory-intensive applications. If offloading behavior is unpredictable, operators may see erratic tail latency and throughput even on identical workloads. By clarifying what makes KV offloading work—or fail—for concurrent agents, the paper provides a practical foundation for building more predictable, memory-efficient LLM serving stacks for agent-centric workloads.