Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Colla-Q方法,通过基于激活熵的比特宽度分配算法平衡MoE模型中各专家性能,提高整体表现并减少对校准数据集依赖。
📝 Abstract
In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
quantization
performance degradation
activation entropy
bit-allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts (MoE)
Quantization
Activation Entropy
Bit Allocation
Minimax Precision Balancing
💼 Related Jobs
No related jobs found.