arXiv:2609.36654v1 Announce Type: cross Abstract: Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosin
Replay the Curvature (arXiv:2609.36654v1) introduces Schur Replay, an algorithm that selects NVFP4 block scales by reproducing GPTQ updates to accurately score reconstruction error. This method, combined with a scalable execution infrastructure, achieves 99.35% and 100.84% question-weighted recovery from BF16 on Qwen3.5-397B and Llama-3.3-70B models. The infrastructure reduces per-layer quantization time by 15.17× over ModelOpt and 23.14× over LLM Compressor while maintaining a low peak memory usage of 35.0 GB per GPU.
The paper addresses a central challenge in low-precision LLM inference: how to make 4-bit weight formats, especially NVFP4, both accurate and practically deployable. NVFP4 represents weights using a very small floating-point codebook, with every 16 E2M1 weights sharing an E4M3 block scale. This block-wise scaling improves local range utilization, but the quality of the resulting model still depends heavily on how those scales—and related rounding or allocation decisions—are chosen. The work frames this as a curvature-aware quantization problem: rather than selecting scales only from per-block amplitude statistics, it uses a “replay the curvature” strategy to estimate which weights or blocks are most sensitive to quantization error.
The key contribution is a scalable quantization method that incorporates second-order sensitivity information without requiring full Hessian storage or expensive per-layer optimization. By replaying representative activations or intermediate signals, the method approximates the quadratic cost of quantization error and uses that signal to guide scale selection in a way that is more aligned with downstream model performance. This makes NVFP4 quantization more robust to outliers, sensitive channels, and the limited representational capacity of 4-bit codebooks, while remaining suitable for large language models.
This matters because LLM inference is often dominated by weight storage and memory traffic, so moving from 16-bit to 4-bit weights can substantially reduce serving cost. However, aggressive quantization can degrade accuracy if the limited codebook is allocated poorly. The paper shows that even in a hardware-relevant, block-wise 4-bit format, curvature information can still be exploited efficiently to improve the accuracy–efficiency tradeoff, making NVFP4 a more attractive option for large-scale inference deployment.