SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
SeqMoE通过预测内存管理和图兼容卸载运行时解决Mixture-of-Experts模型在设备内存限制下的性能问题,提高专家命中率和执行效率。
📝 Abstract
Mixture-of-Experts (MoE) creates a structural advantage for offloading: only a small fraction of activated experts need to reside in device memory, and if they can be loaded in time for computation, offloading can in principle approach full-load performance, where all model weights reside in device memory. Yet translating MoE's structural advantage into practical offloading gains remains challenging. We propose SeqMoE to bridge this gap. To maximize expert hits, we build predictive memory management: (i) Sequence-to-sequence prediction. We are the first to recast expert activation prediction as sequence modeling, enabling accurate multi-step, multi-layer forecasts that provide a long and reliable window for downstream decisions. (ii) Joint prefetch scheduling. We formulate prefetch scheduling as Job Sequencing with Deadlines to maximize expected expert hits and improve bandwidth efficiency. (iii) Forecast-driven caching. Leveraging the recursive nature of sequence modeling, we introduce a probabilistic Belady policy for future-aware eviction. To eliminate execution bottleneck, we develop (iv) Graph-compatible offloading runtime. We derive general runtime principles encompassing compute-transparent expert placement and synchronization-free orchestration disciplines for end-to-end graph capture. With 45% expert residency, SeqMoE averages a 96.97% hit rate and 80.22% of full-load performance, advancing the state of the art in MoE offloading.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
offloading
full-load performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sequence-to-Sequence Prediction
Joint Prefetch Scheduling
Forecast-Driven Caching
Graph-Compatible Offloading Runtime
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zihan Wang
University of Science and Technology of China
Yuqi Wang
Yuqi Wang
The HongKong University of Science and Technology
Oxide SemiconductorDevice ReliabilityProcess Integration
L
Lei Gong
University of Science and Technology of China
C
Cheng Tang
University of Science and Technology of China
Wenqi Lou
Wenqi Lou
University of Science and Technology of China
FPGA AcceleratorAlgorithm-hardware Co-Optimization
Teng Wang
Teng Wang
University of Science and Technology of China
AcceleratorFPGAArchitecture
C
Chao Wang
University of Science and Technology of China
X
Xuehai Zhou
University of Science and Technology of China