arXiv:2609.17943v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a single draft-verify schedule,
ASPIRE is a non-synchronized batched self-speculative decoding framework designed to accelerate long-context LLM inference by allowing individual requests to draft and verify tokens at their own optimal pace, rather than following a synchronized global schedule. It achieves this through three core innovations: a unified mixed forward pass that lets requests in different speculation states coexist in a single batched execution, an online speculation scheduler that dynamically determines per-request verification timing using acceptance-rate estimates and a batch-aware cost model, and a lightweight context-refresh mechanism that updates sparse drafting contexts with minimal overhead.
This approach addresses the memory-bound nature of long-context decoding, where repeated Key-Value (KV) cache reads dominate runtime, by amortizing costly full-attention verification across multiple accepted tokens drafted via efficient sparse attention. Across three models and five reasoning and long-context benchmarks, ASPIRE demonstrates a 1.70–4.58× throughput speedup over standard autoregressive decoding and improves average speedup by approximately 27% over the strongest prior self-speculative baselines.
The paper targets a core bottleneck in long-context LLM serving: attention becomes memory-bound because each decoded step repeatedly reads a large KV cache. Self-speculative decoding is one way to mitigate this by drafting tokens using cheaper sparse attention and then verifying them with full attention, avoiding the need for a separate draft model. However, the authors argue that existing batched self-speculative methods remain overly synchronized: all requests in a batch are forced to follow a shared draft–verify schedule, which can create pipeline bubbles, underutilize hardware, and increase latency when requests differ in context length, draft cost, or acceptance behavior.
The key contribution is ASPIRE, an asynchronous batched self-speculative decoding approach that decouples the drafting and verification stages across requests rather than running the whole batch in lockstep. The central insight is that long-context inference workloads are heterogeneous, so a single global draft–verify cadence is inefficient. By allowing requests to progress asynchronously, ASPIRE can overlap sparse drafting, full-attention verification, and request-level state transitions more effectively. This is primarily a systems-level contribution: it preserves the correctness guarantees of speculative verification while improving scheduling flexibility and hardware utilization in batched serving.
This matters because long-context workloads—such as retrieval-augmented generation, multi-document reasoning, codebase analysis, and agentic inference—are increasingly common, and their performance is often limited by memory-bandwidth-bound decoding rather than raw compute. A method that can improve throughput and reduce latency without requiring a separate draft model is especially practical for production LLM inference stacks. More broadly, ASPIRE points to asynchronous, request-aware execution as an important direction for optimizing speculative decoding in the long-context regime, where traditional batch synchronization can become a significant source of inefficiency.