Attn-QAT: 4-Bit Attention With Quantization-Aware Training

πŸ“… 2026-02-09
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 2
✨ Influential: 1
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of unstable end-to-end 4-bit attention computation, which stems from the extremely limited dynamic range of FP4 and the heavy-tailed distribution of activation valuesβ€”key bottlenecks in low-bit inference. The authors propose Attn-QAT, a quantization-aware training method tailored for attention mechanisms, which enables stable FP4 training without explicit outlier handling by aligning low-precision backward recomputation with corrected gradient assumptions in FlashAttention. They further develop fused Triton kernels for both training and inference. Attn-QAT fully recovers the performance degradation caused by FP4 quantization in diffusion and language models, achieving 1.5Γ— speedup over SageAttention3 and 1.74Γ— over FA4 on an RTX 5090 GPU.
πŸ“ Abstract
Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic range and attention's heavy-tailed activations. This paper presents the first systematic study of 4-bit quantization-aware training (QAT) for attention. We find that"drop-in"QAT, which naively combines an FP4 forward pass with a high-precision Flash Attention (FA)-style backward pass, leads to training instability. We identify two key principles for stable FP4 attention: (1) matching low-precision recomputation of attention scores in the backward pass, and (2) resolving implicit precision assumptions in FA's gradient calculation. Based on these insights, we propose Attn-QAT and implement fused Triton kernels for training as well as FP4 inference kernels. Across diffusion and language models, Attn-QAT recovers the quality drop from FP4 attention without explicit outlier-mitigation heuristics used in prior FP4 attention, and delivers up to a 1.5x speedup on an RTX 5090. Video demos can be found at https://drive.google.com/drive/folders/190F6xbBDUF2kGQYIcXBt3ehSYij5jlim?usp=sharing.
Problem

Research questions and friction points this paper is trying to address.

4-bit attention
quantization-aware training
FP4
dynamic range
heavy-tailed activations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantization-Aware Training
4-bit Attention
FP4
Flash Attention
Low-Precision Training
πŸ”Ž Similar Papers