🤖 AI Summary
To address high energy consumption and latency in deploying Mixture-of-Experts (MoE) models on edge devices—caused by large parameter counts and frequent expert swapping—this work proposes a fine-grained expert caching framework under miss-rate constraints. Our approach comprises three key innovations: (1) Dynamic Bit-Sliced Caching (DBSC), enabling fine-grained, on-demand loading of expert weights; (2) Calibration-free Asymmetric Matryoshka Quantization (AMAT), supporting mixed-precision expert sharing without redundancy; and (3) Predictive Cache Warming (PCW), mitigating cold misses during early decoding stages. Evaluated on DeepSeek-V2-Lite and Qwen1.5-MoE-A2.7B, our method reduces decoding energy consumption by 2.37× and 2.85×, respectively, while improving latency by 1.81× and 1.64×, all with negligible accuracy degradation relative to full-precision baselines.
📝 Abstract
MoE models offer efficient scaling through conditional computation, but their large parameter size and expensive expert offloading make on-device deployment challenging. Existing acceleration techniques such as prefetching or expert clustering often increase energy usage or reduce expert diversity. We present SliceMoE, an energy-efficient MoE inference framework for miss-rate-constrained deployment. SliceMoE introduces Dynamic Bit-Sliced Caching (DBSC), which caches experts at slice-level granularity and assigns precision on demand to expand effective expert capacity. To support mixed-precision experts without memory duplication, we propose Calibration-Free Asymmetric Matryoshka Quantization (AMAT), a truncation-based scheme that maintains compatibility between low-bit and high-bit slices. We further introduce Predictive Cache Warmup (PCW) to reduce early-decode cold misses by reshaping cache contents during prefill. Evaluated on DeepSeek-V2-Lite and Qwen1.5-MoE-A2.7B, SliceMoE reduces decode-stage energy consumption by up to 2.37x and 2.85x, respectively, and improves decode latency by up to 1.81x and 1.64x, while preserving near-high-bit accuracy. These results demonstrate that slice-level caching enables an efficient on-device MoE deployment.