BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决移动设备上MoE模型因DRAM容量有限和数据迁移成本高导致的专家卸载问题,提出BigMoMo方法,通过推测解码优化内存层次结构中的权重重用与计算重叠。
📝 Abstract
Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse computation stalled on weight transfers. Each transfer serves few tokens before execution moves on. We exploit the multi-token verification window of speculative decoding to decouple expert movement from single-token execution, enabling weight reuse, contiguous flash reads, and load-compute overlap. We present \textsc{BigMoMo}, a mobile MoE runtime that exploits this window across the memory hierarchy. It prunes speculative branches and expert activations using acceptance rates, routing impact, and movement cost; reorganizes on-flash experts according to runtime co-loading patterns; and batches ready experts to overlap NPU computation with pending transfers. Across four MoE models and five benchmarks on two mobile platforms, \textsc{BigMoMo} achieves mean decoding speedups of $4.83\times$ over on-demand autoregressive offloading and $1.82\times$ over the best speculative MoE baseline, supporting MoE models up to 30B parameter.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
Efficient Inference
Mobile Devices
Speculative Decoding
DRAM Capacity
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
weight reuse
contiguous flash reads
load-compute overlap
pruning speculative branches
🔎 Similar Papers
No similar papers found.
M
Maoliang Li
School of Computer Science, Peking University, Beijing, China
H
Hailong Zou
School of Computer Science, Peking University, Beijing, China
T
Taohong Han
School of Artificial Intelligence, Tianjin University, Tianjin, China
H
Haoze Chi
School of Electronics Engineering and Computer Science, Peking University, Beijing, China
J
Jiayu Chen
School of Computer Science, Peking University, Beijing, China
Zihao Zheng
Zihao Zheng
Peking University
Machine Learning SystemEdge ComputingComputer ArchitectureEDA
J
Jie Zhang
School of Computer Science, Peking University, Beijing, China
Guojie Luo
Guojie Luo
Peking University
Electronic Design AutomationReconfigurable Architecture
X
Xiang Chen
School of Computer Science, Peking University, Beijing, China