arXiv:2609.15021v1 Announce Type: new Abstract: Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas sharing a 256 GiB LMCache pool. After adopting an existing packed-page patch, we isolate a raw-pointer fallback that omits the

Topological visualization of Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries
Brave API

Based on the provided search context, there is no information regarding the specific paper "Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries" (arXiv:2609.15021v1) or its findings on correctness failures. The available context discusses LMCache and P2P distributed inference systems, which generally achieve correctness by ensuring that stale cache routing results in cache misses rather than incorrect outputs, as the model will simply recompute the prefill if the cache is unavailable.

Key insights from the provided context include:

  • Correctness Mechanism: In distributed KV cache sharing, systems tolerate stale metadata because a routing error leads only to lost optimization (re-prefilling), not semantic violations.
  • Performance Gains: Shared KV caches can reduce time-to-first-token (TTFT) by up to 14x in multi-turn chat workloads by eliminating redundant prefill computation across data-parallel ranks.
  • Workload Suitability: Significant benefits are observed in multi-turn chat, agents, and long shared system prompts, whereas unique one-shot prompts see minimal gains.
  • Architecture: Solutions like LMCache move KV caches out of GPU memory into a tiered hierarchy (CPU, disk, remote storage) to enable sharing across processes and nodes.
Generated 19d ago
Open-Weights Reasoning

The paper examines shared host-memory KV caching as a mechanism for reducing redundant prefill in replicated large-model inference. The setup is concrete: two single-GPU vLLM replicas serving a 27B model share a 256 GiB LMCache pool, so that when a request is rerouted between replicas, the KV state for its prompt prefix can in principle be reused instead of recomputed. The authors treat this not merely as a performance optimization, but as a systems-correctness problem: shared KV caching is only useful if the cached state is transferred correctly and if the request pattern actually preserves enough prefix locality to make the reuse worthwhile.

A key contribution is the isolation of a correctness failure in an existing packed-page patch. The authors find that a raw-pointer fallback path can omit the state or metadata needed to make cached KV pages portable across replicas, producing incorrect state transfer rather than a safe miss. In other words, the system may silently use a malformed or incomplete cache entry instead of falling back to recomputation. The paper then maps the performance boundaries of the approach: even when state transfer is correct, the benefit of shared KV caching depends heavily on whether the receiving replica actually has the relevant prefix cached, how much of the prefix is reusable, and how host-memory capacity and transfer overhead interact with the 27B model’s serving workload.

This matters because shared KV caching is increasingly attractive in multi-replica LLM serving, where load balancing, autoscaling, or request migration can otherwise force expensive repeated prefill. The work provides a cautionary but practical result: cache sharing must be validated for correctness first, and its performance gains should not be assumed from cache size alone. For practitioners deploying vLLM/LMCache-style systems, the paper highlights both a concrete bug class in state transfer and the operational conditions under which shared KV caching can actually improve serving efficiency.

Generated 19d ago
Sources