Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过使用八个基于Nsight Compute报告的视图,分析了在Nvidia Hopper GPU上LLM推理过程中利用率低下的具体原因及机制。
📝 Abstract
A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is wrong, but that it collapses several different mechanisms into one number. This is most severe during decode, where each request contributes only one new token and dense projection GEMMs become small-row matrix multiplications. On Hopper, the bfloat16 GMMA path executes these operations in fixed 64-row matrix fragments, so small-batch decode can fill only a small fraction of each fragment with real token rows. In this paper, we profile vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length and batch size. We replace the usual single utilization number with eight counter-validated views derived from raw Nsight Compute reports, each pinned to an NCU counter or explicit formula. Together, these views map utilization gaps to concrete mechanisms - fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection - across four production models and six per-layer kernel roles.
Problem

Research questions and friction points this paper is trying to address.

GPU Utilization
LLM Inference
Hopper
Decode
Fragment Fill
Innovation

Methods, ideas, or system contributions that make the work stand out.

fragment fill
occupancy limits
stall signatures
wave quantization
kernel selection
🔎 Similar Papers
No similar papers found.