arXiv:2609.19499v1 Announce Type: cross Abstract: Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produc
Sample Count Is Not Enough is a September 2026 paper by Mobina Kashaniyan and Ali Jannesari (arXiv:2609.19499) demonstrating that how candidate responses are grouped into generation calls significantly impacts the energy and performance of LLM test-time scaling, not just how many candidates are generated.
The study finds that for a fixed candidate count (e.g., $N=8$), splitting generation across more sequential calls with smaller batch sizes consistently increases latency, GPU-hours, and gross GPU-device energy while reducing throughput compared to fewer, larger batched calls.
Key findings include: Efficiency Rule: When candidates are independent and memory allows, fewer generation calls with larger batch sizes are more efficient; for example, moving from 8 serial calls to 1 batched call can save approximately 1.37 MWh of energy and 10,401 GPU-hours per million queries. Reporting Gap: The authors argue that candidate count $N$ alone is insufficient to describe system cost; evaluations must report generation schedule, latency, throughput, and energy metrics to ensure reproducibility. * Scope: Experiments on Phi-3-mini and Qwen2.5-1.5B models using GSM8K and SciQ datasets on NVIDIA A100 and V100 GPUs confirmed these trends across different hardware and output lengths.
This material examines test-time scaling for large language model reasoning, where multiple candidate responses are generated and then combined to improve final-answer quality. It challenges the common convention of describing an inference budget only by the number of generated candidates, N. The central argument is that N is an incomplete budget descriptor: it specifies how many candidates are produced, but not how they are produced. For a fixed candidate count, different execution strategies—such as variations in scheduling, batching, parallelism, decoding order, stopping behavior, or candidate selection—can lead to materially different runtime, hardware utilization, energy consumption, and even final reasoning performance.
A key contribution is reframing the evaluation of test-time scaling around candidate-generation strategy rather than sample count alone. The work highlights that two systems with the same nominal candidate budget can occupy different points in the accuracy–energy tradeoff space. In other words, the marginal value of additional candidates depends on the generation policy, and energy cost is not determined solely by the total amount of sampled text. This provides a more resource-aware view of test-time scaling, emphasizing that performance and efficiency should be assessed jointly rather than treating “more samples” as a neutral or fully specified resource choice.
This matters because test-time scaling is increasingly used to push LLMs on difficult reasoning tasks, but deployment decisions are constrained by latency, cost, and energy. By showing that candidate-generation strategy is a first-class design variable, the material offers a more complete framework for benchmarking and engineering scalable inference. It suggests that optimizing how candidates are generated can be as important as increasing their number, helping practitioners target the right accuracy–efficiency frontier for real-world use.