ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the rounding ambiguity in calibration-free low-bit quantization of large language models, where standard round-to-nearest (RTN) suffers from instability due to weights clustering near quantization bin boundaries. To resolve this, the authors propose ReRound, a novel method that leverages a conditional diffusion model to reconstruct low-bit weights as a guiding signal. ReRound dynamically selects between reconstruction-guided rounding and conventional RTN based on a tolerance metric and further refines the quantized weights via singular value matching to preserve structural fidelity. Evaluated at 3- and 4-bit precisions, ReRound significantly outperforms existing calibration-free approaches, achieving performance on par with calibration-dependent methods while introducing no overhead during inference. This establishes a new paradigm for calibration-free quantization grounded in generative reconstruction and structure-aware optimization.
📝 Abstract
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusion model to produce continuous reconstructions of low-bit weights for the LLM. These reconstructed weights act as a guidance signal to disambiguate the rounding direction of weights located close to interval midpoints. To integrate this reconstruction-guided rounding with conventional RTN, ReRound introduces a tolerance metric measuring how far the quantized weight (not the final quantized integer) is away from the midpoint: quantized weights within a tolerance region around midpoints are quantized using diffusion-based reconstructions, whereas weights closer to quantization boundaries are quantized with RTN. By sweeping the tolerance parameter, ReRound generates multiple candidate quantized integer weight matrices and selects the de-quantized weight matrix candidate whose leading singular values most closely match those of the original full-precision weights. This selected candidate determines the tolerance parameter ReRound uses. ReRound is particularly effective for smaller LLMs. Across a range of such models, it consistently outperforms standard RTN for 3-bit and 4-bit weight quantization. ReRound achieves superior accuracy compared to an extensive set of calibration-free methods, remains competitive with calibration-dependent approaches, and operates entirely offline, introducing no additional overhead during low-bit inference. The ReRound strategy represents a new approach for low-bit quantization. The method applies to AI models beyond LLMs. This paper focuses on its applications to small LLMs.
Problem

Research questions and friction points this paper is trying to address.

midpoint ambiguity
LLM quantization
post-training quantization
rounding error
calibration-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reconstructive Rounding
midpoint ambiguity
diffusion-based reconstruction
calibration-free quantization
post-training quantization
🔎 Similar Papers
No similar papers found.