RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying extremely large Mixture-of-Experts (MoE) language models (26–120B parameters) on consumer-grade hardware, where excessive memory demands from model weights, KV caches, and expert sublayers hinder practical inference. To overcome this, the authors propose a three-axis compression framework comprising architecture-aware mixed-precision quantization (2/4/8-bit), an LRU-driven expert offloading mechanism, and IsoQuant—a novel KV cache compression method that combines Walsh–Hadamard transforms with SO(4) rotations to enable efficient low-rank isotropic quantization. IsoQuant further incorporates end-to-end fused GPU kernels for direct 3-bit tensor processing. Experiments demonstrate that the approach enables Gemma 4-26B-A4B and Qwen3-30B-A3B to run on 16GB GPUs and Nemotron-H 120B on 32GB GPUs, achieving inference speeds of 9–19 tokens per second, perplexity degradation no worse than +0.0012, and perfect 100% retrieval accuracy over 32K-context evaluations.
📝 Abstract
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring $O(d \log d)$ operations and 256 stored parameters per head versus $O(d^2)$ and 16{,}384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation ($Δ$PPL $\leq +0.0012$) and 100\% retrieval accuracy at 32K context.
Problem

Research questions and friction points this paper is trying to address.

mixture-of-experts
memory constraint
consumer hardware
large language models
KV cache
Innovation

Methods, ideas, or system contributions that make the work stand out.

RotaryQuant
IsoQuant
mixed-precision quantization
KV cache compression
MoE models
A
Anthony Lui
Cognizant AI & Analytics, London, UB8 3PH, UK
M
Mohamed Elsaied
Cognizant AI Lab, Responsible AI Office, San Francisco, USA; Southern Methodist University (SMU), Engineering Management, Lyle School of Engineering, Dallas, TX, USA
N
N. P. Savani
Cognizant AI Lab, Responsible AI Office, San Francisco, USA; University of Maryland, Baltimore County, Goddard Planetary Heliophysics Institute, 5523 Research Park Drive, Baltimore, MD 21228, USA