arXiv:2609.17573v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) are emerging as a promising generative paradigm that complements autoregressive decoding. In long-context settings, KV cache bloat and offloading transfer overhead have become primary bottlenecks in inference systems. Meanwhile, the periodic full-sequence recomputation and localized token updates in dLLMs ma

Topological visualization of GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
Brave API

GroupKV is a hierarchical KV cache management system designed to address the bottlenecks of KV cache bloat and offloading transfer overhead in long-context Diffusion Large Language Models (dLLMs). Unlike autoregressive models that use an append-only cache, dLLMs require periodic full-sequence recomputation and localized token updates, creating a dynamic KV lifecycle that complicates traditional caching strategies.

To solve this, GroupKV utilizes a coarse-to-fine sparse selection approach, partitioning context into contiguous groups and exploiting intra-block spatial locality and cross-layer consistency to prioritize retrieval. The system incorporates predictive prefetching to overlap data transfer with computation, a staleness correction mechanism to maintain coherence during dynamic updates, and streaming prefill to reduce peak GPU memory usage.

Experimental results demonstrate that GroupKV extends the maximum serviceable context length by up to 48.0× under constrained GPU memory and improves end-to-end inference performance by up to 3.73× in offload-based settings, while maintaining competitive accuracy on benchmarks like LongBench and RULER. This makes GroupKV particularly effective for resource-constrained long-context inference where existing methods fail due to memory limits.

Generated 17d ago
Open-Weights Reasoning

The paper addresses a key systems bottleneck in long-context inference for diffusion large language models (dLLMs). Unlike autoregressive decoding, dLLMs refine a sequence through iterative denoising or diffusion steps, which can trigger periodic full-sequence recomputation even when only a subset of tokens changes. In long-context settings, this behavior compounds the usual KV-cache memory pressure: the cache grows with context length, and offloading cache state to host memory or storage introduces substantial transfer overhead. GroupKV is introduced as a hierarchical KV-cache management approach designed to reduce these costs by exploiting the structure of dLLM inference.

GroupKV organizes KV cache state into a hierarchy of token groups rather than managing individual token cache entries in a flat manner. This grouping enables more efficient decisions about which cache regions to keep in fast memory, which to offload, and which can be reused or selectively recomputed across diffusion steps. Because dLLM updates are often localized and many context regions remain relatively stable during a generation episode, the system can avoid redundant recomputation and minimize the amount of KV state moved between compute and storage tiers. The central insight is that cache management for dLLMs should be aligned with the model’s iterative update pattern, not just with static notions of token importance.

The work matters because it targets one of the practical limits preventing dLLMs from being competitive with autoregressive systems on long-context workloads. By reducing memory footprint, offloading overhead, and unnecessary recomputation, GroupKV can improve latency and throughput for long-document generation, iterative refinement, agentic workflows, and other settings where dLLMs may offer complementary strengths. More broadly, it points toward a systems-level design principle for diffusion LLM inference: efficient serving requires cache policies that understand both the memory hierarchy and the temporal dynamics of diffusion decoding.

Generated 17d ago
Sources