arXiv:2609.28870v2 Announce Type: replace Abstract: Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a la
The paper When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse (arXiv:2609.28870v2) argues that sophisticated eviction algorithms designed for traditional caches provide little benefit over simple LRU for LLM prefix caching.
This inefficiency stems from the structural nature of prefix reuse, which is dominated by the regular pacing of active sessions, making recency an unusually predictive signal while frequency proves ineffective.
Effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity.
The paper studies prefix caching in LLM serving, where repeated requests share long context prefixes and KV-cache reuse can avoid redundant prefill computation. It focuses on long-running, agentic applications whose contexts grow over time, a regime in which cache replacement decisions are especially consequential because old prefixes, tool outputs, and evolving task state compete for limited accelerator memory. The authors analyze production traces from two companies and benchmark 14 eviction algorithms across both HBM-constrained and large memory-pool configurations, positioning the work as an empirical investigation of how prefix-cache replacement actually behaves under realistic multi-turn workloads.
Its main contribution is a systematic comparison of more sophisticated cache-replacement policies against simpler or more direct alternatives for LLM prefix reuse. Rather than assuming that advanced eviction heuristics—such as those inspired by general-purpose caching or designed to predict future reuse—will automatically improve prefix-cache hit rates, the paper evaluates them under the access patterns induced by agentic loops, growing prompts, and heterogeneous context lifetimes. The result is a more nuanced view of when eviction sophistication helps and when it fails, highlighting that the best replacement strategy can depend on memory capacity, workload locality, and the degree to which prefixes are reused across turns or sessions.
This matters because prefix caching is one of the main levers for reducing prefill latency and GPU cost in modern LLM systems, especially as applications become more stateful and agent-like. By grounding the analysis in production traces and multiple memory regimes, the paper provides practical guidance for serving-system designers: cache replacement should be tuned to the structure of LLM prefix reuse rather than adopted from generic caching heuristics without validation. It also identifies an important open problem in LLM infrastructure—making prefix caches efficient under long-lived, evolving contexts—where better replacement policies can directly improve throughput, tail latency, and serving economics.