LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization

📅 2026-04-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the significant degradation in generation quality caused by 4-bit quantization in large diffusion Transformers. Existing low-rank compensation methods rely on high-precision auxiliary branches and data-intensive calibration, limiting their practicality. To overcome these constraints, we propose LoRaQ, a calibration-free low-rank approximation quantization method that explicitly optimizes quantization error compensation. LoRaQ enables the low-rank branch itself to be quantized to sub-16-bit precisions—such as W4A8, W6A6, and W8A8—establishing the first fully sub-16-bit inference pipeline for diffusion models. Without requiring high-precision branches or calibration data, LoRaQ achieves superior performance over state-of-the-art methods on Pixart-Σ and SANA under identical memory budgets and demonstrates the effectiveness of diverse mixed-precision configurations.

Technology Category

Application Category

📝 Abstract
Post-training quantization (PTQ) is essential for deploying large diffusion transformers on resource-constrained hardware, but aggressive 4-bit quantization significantly degrades generative performance. Low-rank approximation methods have emerged as a promising solution by appending auxiliary linear branches to restore performance. However, current state-of-the-art approaches assume these branches must retain high precision (W16A16) and rely on heavy, data-dependent calibration for initialization. We challenge both limitations with LoRaQ (Low-Rank Approximated Quantization), a simple, data-free calibration approach that optimizes quantization error compensation. By overcoming the need for high-precision branches, LoRaQ enables the first fully sub-16 bit pipeline, allowing the low-rank branch itself to be quantized. We demonstrate that, at equal memory overhead, LoRaQ outperforms the state-of-the-art methods in their native implementations on Pixart-$Σ$ and SANA. We also analyze mixed-precision configurations, showing that setups such as W8A8, W6A6, and W4A8 for the low-rank branch, alongside a W4 main layer, yield superior results while maintaining a fully quantized architecture compatible with modern mixed-precision hardware.
Problem

Research questions and friction points this paper is trying to address.

post-training quantization
low-rank approximation
4-bit quantization
diffusion transformers
quantization error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Rank Approximation
4-bit Quantization
Post-Training Quantization
Data-Free Calibration
Mixed-Precision
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yann Bouquet
Research and Advancement Development Team, Advanced Micro Devices, Inc, USA; Computer Vision Laboratory, Ecole Polytechnique Fédérale de Lausanne, Lausanne, Switzerland
Alireza Khodamoradi
Alireza Khodamoradi
Researcher at AMD
AI and HPC
S
Sophie Yáng Shen
Research and Advancement Development Team, Advanced Micro Devices, Inc, USA
Kristof Denolf
Kristof Denolf
Principal Engineer, Xilinx
Cost Efficient Vision Processing
Mathieu Salzmann
Mathieu Salzmann
EPFL
Computer visionmachine learning