Quantization Degradation in Large Language Models: A Signal-Noise Perspective

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the causes of performance degradation in large language models after weight quantization and examines how this degradation varies across bit widths, quantization methods, model scales, and tasks. Innovatively adopting a signal-to-noise perspective, the work decomposes quantization error into its intra-layer generation mechanisms and inter-layer propagation dynamics through signal-to-noise ratio (SNR) analysis, thereby attributing quantization-induced performance lossβ€” for the first timeβ€”to the coupling between the SNR characteristics of error sources and their propagation behavior. Experimental results demonstrate that 4-bit quantization generally preserves model performance, 2-bit quantization consistently leads to degradation, and 3-bit performance is highly dependent on the specific task, quantization method, and model size. Notably, larger models exhibit greater robustness due to their inherent attenuation of error amplification across layers.
πŸ“ Abstract
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
Problem

Research questions and friction points this paper is trying to address.

quantization degradation
large language models
post-training quantization
signal-to-noise ratio
model compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

quantization degradation
signal-to-noise ratio
post-training quantization
error propagation
large language models
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
C
Chenxi Zhou
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China
Pengfei Cao
Pengfei Cao
Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelsInformation Extraction
J
Jinyu Ye
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
B
Bohan Yu
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China
H
Haida Yu
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
J
Jiang Li
College of Computer Science, Inner Mongolia University, Hohhot, China
Jun Zhao
Jun Zhao
School of Marine Sciences, Sun Yat-sen University
ocean opticsremote sensingnumerical modeling
K
Kang Liu
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China