REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型在资源受限环境下的部署问题,REAL-Q通过动态梯度下降和细粒度修正方法优化了后训练量化过程,减少了端到端KL散度。
📝 Abstract
Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.
Problem

Research questions and friction points this paper is trying to address.

Post-training Quantization
Large Language Models
Information Misalignment
Error Propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

REAL-Q
E2E-loss Aligned
Dynamic Block-wise Gradient Descent
Sliding Window Mechanism
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Q
Qian Zhang
Peking University
Y
Yaoming Li
Peking University
Z
Zhewen Tan
Peking University
Yanshu Wang
Yanshu Wang
Tsinghua University
Computer&Economy
Heng Lu
Heng Lu
Microsoft
speech synthesisspeech recognitionNLPdeep learningmachine learning
Kun Su
Kun Su
Google Research
Multimodal LearningAudio/Music GenerationRecommendation system
Z
Zongwei Lv
Peking University
W
Wenhan Yu
Peking University
Y
Yongge Ma
Peking University
Y
Yinjun Han
ZTE Corporation
R
Ruikuang Liu
ZTE Corporation
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookupHash Table