A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number. This is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multipli

Topological visualization of Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
Brave API

The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number, specifically hiding how much useful work is actually being done.

This issue is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multiplications. A single SM utilization percentage can make a workload look compute-saturated while masking bottlenecks such as:

  • Fragment fill limits: Hopper’s BF16 GMMA m64 fragment floor forces small-batch matrix multiplies to pad most rows with zeros.
  • Wave quantization: Launch grid inefficiencies leave SM slots idle in the trailing wave.
  • Resource-limited occupancy: Kernels are capped by register and shared memory limits rather than hardware concurrency.
  • Memory stalls: Significant time is spent waiting for data (long_scoreboard stalls) rather than computing.

Replacing the single scalar with eight counter-validated views reveals that decode underutilization is a composition of these distinct mechanisms, each requiring a different fix rather than simple library tile retuning.

Generated 19d ago
Open-Weights Reasoning

This material examines what “GPU utilization” actually measures for LLM inference on Nvidia Hopper, arguing that a single SM utilization percentage is an overly coarse headline metric. It is not necessarily wrong, but it can be misleading because it aggregates several distinct hardware behaviors into one number: SMs being active, Tensor Cores issuing useful matrix operations, memory subsystems being busy, and instruction pipelines being stalled or underfed. The paper’s central concern is that a workload can appear compute-saturated while doing relatively little useful arithmetic, especially during the decode phase of inference, where each request generates only one new token per step and many dense projection GEMMs degenerate into small-row matrix multiplications.

Its key insight is that high SM occupancy or high SM active cycles can coexist with low effective compute throughput. In decode, the relevant GEMM shapes often have a very small row dimension, which limits arithmetic intensity and can leave Tensor Cores underutilized even when SMs look busy. The paper therefore dissects utilization into more informative components—such as Tensor Core activity, memory bandwidth pressure, kernel shape effects, occupancy, instruction mix, and stall behavior—rather than treating “SM utilization” as a single performance verdict. This reframing helps explain why inference engines may show strong GPU activity while still leaving substantial headroom for better batching, sequence packing, kernel selection, quantization, or memory-system optimization.

This matters because LLM serving is increasingly governed by decode efficiency, where raw FLOPs are less predictive of throughput than the ability to keep memory and compute pipelines effectively fed. A misleadingly high SM utilization number can make a deployment look saturated and discourage further optimization, or cause benchmark comparisons to overstate the efficiency of an inference stack. By providing a more granular diagnostic lens for Hopper-class GPUs, the work gives practitioners a better basis for capacity planning, profiling, and tuning inference systems around the workloads that actually dominate production latency and cost.

Generated 19d ago
Sources