ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
This work addresses the significant storage overhead incurred by retaining the language model head (LM-head) in high-precision formats after weight quantization, as direct quantization severely distorts the logit distribution. To overcome this challenge, the authors propose ARCHead, a novel LM-head compression method that introduces, for the first time, an activation-aware metric combined with low-rank decomposition, grouped INT4 residual quantization, and a low-rank correction mechanism. This approach achieves substantial storage reduction with nearly negligible performance degradation. Evaluated on Qwen3-8B-Base, ARCHead requires only 25.6% of the storage of a BF16 head while achieving a relative perplexity of 1.007. When replacing the original BF16 head, it incurs a minimal cross-entropy increase of merely 0.006–0.007 and reduces throughput by less than 2%.