arXiv:2609.39131v1 Announce Type: cross Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increas
Characterizing High Bandwidth Flash for LLM Serving (arXiv:2609.39131v1) evaluates the use of High-Bandwidth Flash (HBF) to overcome memory bottlenecks in large language model serving. The study introduces an HBM–HBF–host hierarchical storage system and buffered cache-aware scheduling to optimize data placement and admission control.
Key findings indicate that HBF-augmented systems can reduce workload completion time by 46–56% and model energy consumption by 14–24% relative to HBM-only baselines. Crucially, buffering admission by just 10% headroom extends the estimated HBF write lifetime from 4.8 to 14.8 years, demonstrating that coordinated scheduling makes HBF a viable, durable solution for expanding accelerator memory capacity.
This material examines high-bandwidth flash (HBF) as a memory option for large language model serving, where the primary constraints are increasingly the capacity and bandwidth needed to hold model weights and KV caches near accelerators. It frames the problem around the fact that modern LLM workloads—especially long-context and agentic workloads—create repeated, stateful memory traffic: prompts, multi-turn agent histories, and growing KV caches continually expand the amount of data that must be read, written, and retained during inference. The paper’s central contribution is a characterization of how such serving workloads stress memory systems, and how HBF’s capacity, bandwidth, latency, and endurance tradeoffs compare with conventional DRAM/HBM-centric designs.
A key insight is that LLM serving does not impose a single uniform memory demand. Prefill tends to be bandwidth- and compute-intensive, decode is more sensitive to latency and the availability of hot KV state, and agentic loops amplify the KV footprint by repeatedly accessing or extending long contexts. HBF is therefore best viewed not simply as a slower replacement for DRAM, but as a tiering and capacity-scaling mechanism: it can potentially absorb cold weights, long KV caches, or agent state, while faster memory handles the most latency-critical data. The work contributes a design-level understanding of what bandwidth, capacity, and access-pattern requirements are needed for HBF to be effective in serving stacks, and where system techniques such as KV cache management, offloading, prefetching, and memory hierarchy placement become essential.
This matters because economically scalable LLM serving will likely require moving beyond GPU-local HBM as the sole source of capacity. If high-bandwidth flash can support large model weights, long contexts, and multi-turn agent state at lower cost per gigabyte, it could enable higher concurrency, longer sessions, and larger models without proportionally increasing accelerator memory cost. The characterization provides a practical basis for memory-system architects, inference engineers, and hardware designers to reason about when flash-based memory is sufficient, where it becomes a bottleneck, and what co-design choices are needed to meet real serving latency and throughput targets.