Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction algorithms operate at a token granularity finer than block-level management. This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memo
Based on the provided search context, there is no information regarding "vToken" or "Token-Level Virtualization for Reclaimable KV Caches." The context exclusively discusses PagedAttention and vLLM, which utilize block-level (not token-level) virtualization, dividing the KV cache into fixed-size blocks (typically 16 or 32 tokens) to manage memory.
Key points from the available context regarding current state-of-the-art methods include: PagedAttention inspired by OS virtual memory manages the KV cache in fixed-size blocks, mapping logical blocks to non-contiguous physical memory. This approach reduces memory fragmentation from 60–80% (in contiguous pre-allocation systems) to under 4%. It enables 2x to 4x higher throughput by allowing larger batch sizes and eliminating both internal and external fragmentation. Current optimizations focus on block-level management; token-level memory management is mentioned only as a direction for future follow-up work, not as an implemented system like vToken.
vToken addresses a subtle but important inefficiency in modern LLM inference serving: even when KV-cache memory is managed with block-based allocators such as PagedAttention, token-granular eviction and reuse policies can leave partially used blocks unable to be reclaimed. The paper argues that this creates intra-block fragmentation, where the physical KV memory is allocated, but only some of the tokens inside a block are still useful or eligible for reuse. As a result, systems may reserve a substantial amount of KV memory that cannot be safely reused, reducing effective cache capacity and limiting batch size or context length.
The central contribution is a token-level virtualization mechanism for reclaimable KV caches. Rather than forcing the serving system to choose between coarse block allocation and fine-grained token eviction, vToken separates the logical management of individual KV tokens from the physical organization of KV memory into blocks. This abstraction allows the system to track which tokens within a block are reclaimable, expired, or still needed, and to reclaim or repurpose memory at token granularity while still benefiting from block-based allocation. In effect, it makes partially occupied blocks usable again instead of treating them as fully consumed.
This matters because KV-cache memory is one of the primary constraints in high-throughput LLM serving. By reducing intra-block waste, token-level virtualization can increase the effective amount of usable KV cache, improve memory utilization under eviction-heavy workloads, and enable more aggressive caching or reuse policies without requiring a redesign of the underlying block allocator. For serving systems that combine PagedAttention-style memory management with token-level KV eviction, vToken offers a practical path to better capacity, higher concurrency, and lower memory overhead.