MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers

๐Ÿ“… 2026-07-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing post-training quantization methods for Vision Transformers employ uniform bit-widths across all components, disregarding their heterogeneous sensitivity to quantization and thereby incurring substantial accuracy degradation. This work proposes MixFrag, a novel framework that introduces, for the first time, a KL divergenceโ€“based metric to quantify the quantization fragility of each layer. By leveraging a small calibration set to assess output distribution shifts, MixFrag formulates mixed-precision bit allocation as a multiple-choice knapsack problem, enabling adaptive optimization under a given bit budget. The method achieves state-of-the-art performance on ImageNet-1K classification as well as COCO detection and segmentation tasks, yielding improvements of up to 9.6 AP over prior best approaches under MP3/MP3 settings.
๐Ÿ“ Abstract
Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However, existing PTQ methods typically employ uniform bit-widths across transformer components, overlooking their heterogeneous sensitivity to quantization and leading to inefficient precision allocation. In this paper, we propose {MixFrag, a fragility-guided mixed-precision PTQ framework for Vision Transformers. MixFrag first estimates component-level quantization fragility by measuring the Kullback--Leibler (KL) divergence between full-precision and isolated quantized output distributions using a small calibration set. It then formulates bit allocation as a Multiple-Choice Knapsack Problem (MCKP), enabling adaptive layer-wise precision assignment under a target bit budget. Extensive experiments on ImageNet-1K across multiple Vision Transformer architectures demonstrate that MixFrag achieves competitive classification performance under practical mixed-precision settings. Furthermore, evaluations on COCO object detection and instance segmentation show that MixFrag achieves state-of-the-art performance among existing mixed-precision PTQ methods, improving the previous best method by up to 9.6 AP under the challenging MP3/MP3 setting. Additional analyses validate the proposed fragility metric and demonstrate its strong correlation with the learned bit allocation. These results establish MixFrag as an effective framework for mixed-precision post-training quantization of Vision Transformers.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
post-training quantization
mixed-precision
quantization fragility
bit-width allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

mixed-precision quantization
post-training quantization
Vision Transformers
quantization fragility
Multiple-Choice Knapsack Problem
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Md. Mehrab Hossain Opi
Department of Computer Science and Engineering, Khulna University of Engineering & Technology (KUET), Khulna 9203, Bangladesh
R
Robiul Islam Ryad
Department of Computer Science and Engineering, Khulna University of Engineering & Technology (KUET), Khulna 9203, Bangladesh
M
Md. Umar Faruk
Department of Computer Science and Engineering, Khulna University of Engineering & Technology (KUET), Khulna 9203, Bangladesh