arXiv:2609.17652v1 Announce Type: cross Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which each query decides how many bits of each key channel to read. The 4-bit
Fathom is a sparse decoding method designed for scenarios where KV caches and their ranking indices reside in host memory rather than GPU memory. It optimizes bandwidth-bound decode steps by storing the 4-bit Key cache as channel-major bit planes, allowing each query to perform per-query read depth selection via reverse water-filling over channel importance.
In the target regime (e.g., 1 million tokens on Qwen3-8B), Fathom achieves a 1.67x faster GPU time than 136-bit scans (like Double Sparsity, Loki, and SparQ r=32) and moves 18% fewer bytes than SparQ r=16 at equal GPU time. It reaches the step agreement accuracy of the most accurate 136-bit scan using only 92 bits, though it provides no speedup if the index is already resident in GPU High Bandwidth Memory (HBM).
Fathom targets a memory-bandwidth bottleneck in long-context sparse decoding. In agentic workloads with very long sessions—potentially on the order of a million tokens—and many concurrent sessions, the KV cache and the index used to rank candidate KV entries typically reside in host memory. For sparse decoding, each step must often scan a large set of keys to select the top-k most relevant cache entries. Fathom identifies this ranking scan, rather than the attention computation itself, as the dominant traffic source that can bound decoding throughput when the cache is offloaded.
The key idea is to make the key-ranking scan precision-aware and query-dependent. Instead of reading every key at a fixed precision or width, Fathom lets each query determine how many bits of each key channel to read during the scan. This per-query read-depth mechanism allows the system to avoid unnecessary high-precision reads when lower-bit information is sufficient for ranking, while still preserving the accuracy of the top-k selection. The approach is especially relevant because the cost of scanning all candidate keys grows with context length and concurrency, so even modest reductions in per-key read width can translate into meaningful bandwidth savings.
The contribution matters because it addresses a practical scaling problem for offloaded KV caches. As LLM systems move toward longer contexts, multi-agent sessions, and many resident conversations, host-memory KV caches become attractive for capacity but are limited by memory traffic. Fathom reframes sparse decoding as a read-depth optimization problem, coupling key ranking with adaptive precision. If effective, this can improve decoding efficiency for long-context inference without requiring the full KV cache to remain in fast device memory, making sparse attention more practical for memory-constrained, high-concurrency serving environments.