Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对量化语言模型中的不确定性保留问题,提出了一种基于目标的校准数据选择方法Doubt-Preserving Quantization (DPQ),以适应不同部署需求。
📝 Abstract
Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abstention, is rarely treated as a primary objective. We frame calibration-data selection for quantization as a target-dependent uncertainty-preservation problem. Different deployments emphasize different regions of the input distribution, yet prior work mainly optimizes accuracy-oriented compression metrics or adjusts scores after quantization. We formalize this goal with distributional and boundary preservation risks, and provide a simple mixture-mismatch argument explaining why no single calibration recipe should be expected to fit all targets. We introduce Doubt-Preserving Quantization (DPQ), a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors. Across 8 language models, 9 NLP benchmarks, and 22 comparison methods, the leading fixed recipe changes with the preservation target: DPQ-r75 leads on SQuAD2 answerability-boundary preservation, while milder or single-signal variants, including DPQ-r50, confidence-only, and entropy-only, better preserve broad multiple-choice QA behavior. These results show that calibration data should be selected for the specific full-precision score behavior a deployment needs to preserve, rather than treated as a fixed quantization detail.
Problem

Research questions and friction points this paper is trying to address.

Quantization
Uncertainty Preservation
Calibration Data Selection
Language Models
Deployment Specific
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target-Aware Calibration
Uncertainty Preservation
Doubt-Preserving Quantization
Distributional Risk
Boundary Preservation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zhen Yang
Yale University
S
Sizai Hou
The Hong Kong University of Science and Technology
K
Kaiwen Zheng
The Hong Kong University of Science and Technology (Guangzhou)
Yaofang Liu
Yaofang Liu
City University of Hong Kong
Diffusion ModelsVideo GenerationImage Processing
L
Liang He
Shanghai Institute of Optics and Fine Mechanics
Yixuan Chen
Yixuan Chen
Oxford Suzhou Center for Advanced Research
DisentanglementVision-Language ModelAI for Medical
Kangning Cui
Kangning Cui
Research Assistant Professor of Computer Science, Wake Forest University
Applied MathematicsComputational SustainabilityMedical Imaging