HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决MoE模型在3D近存处理架构上的高效部署问题,HDA-MoE通过混合并行部署与动态自适应调度方法减少通信开销并提高计算利用率。
📝 Abstract
Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However, this efficiency comes at the expense of increased memory capacity and bandwidth demands. Recent 3D Near-Memory Processing (NMP) architectures, which vertically integrate memory and compute through hybrid bonding, provide high internal bandwidth and energy efficiency, making them attractive for accelerating MoE inference. Nevertheless, the distributed memory and compute organization of NMP systems introduces new challenges for mapping MoE workloads. Existing parallelization strategies, such as Tensor Parallelism (TP) and Expert Parallelism (EP), suffer from either high communication costs or unbalanced computation utilization, leading to inferior efficiency. In addition, the dynamic routing behavior of MoE models further complicates efficient deployment. To address these challenges, we present HDA-MoE, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling. HDA-MoE integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation utilization. Experimental results show that HDA-MoE achieves a speedup of 1.1x--3.4x over TP, 1.1x--1.5x over EP, 1.1x--3.7x over the Hybrid TP-EP compute-balanced baseline, and 1.1x--1.3x over HD-MoE. Source code is available at https://github.com/PKU-SEC-Lab/HDA-MoE-TCAD26.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
3D Near-Memory Processing
communication cost
computation utilization
dynamic routing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Parallelism
Dynamic and Adaptive Scheduling
3D Near-Memory Processing
Mixture-of-Experts
💼 Related Jobs
No related jobs found.
Haochen Huang
Haochen Huang
University of California San Diego
system/software reliabilitysecurity
Shuzhang Zhong
Shuzhang Zhong
Peking University
Machine Learning System
S
Shengxuan Qiu
Institute for Artificial Intelligence, Peking University, Beijing, China; School of Integrated Circuits, Peking University, Beijing, China; School of Electronics Engineering and Computer Science, Peking University, Beijing, China
Z
Zhe Zhang
Alibaba DAMO Academy, Beijing, China; DAMO Academy, Alibaba Group, Beijing Hupan Lab, Hangzhou
Shuangchen Li
Shuangchen Li
Research Scientist, DAMO Academy, Alibaba Group
Computer ArchitectureElectronic Design Automation
Cong Li
Cong Li
School of Integrated Circuits, Peking University
Computer ArchitectureNear Data Processing
Dimin Niu
Dimin Niu
Computing Technology Lab, Alibaba DAMO Academy
Computer ArchitectureMemory SystemsProcessing-in-MemoryDeep Learning
H
Hongzhong Zheng
Alibaba DAMO Academy, Beijing, China; DAMO Academy, Alibaba Group, Beijing Hupan Lab, Hangzhou
Guangyu Sun
Guangyu Sun
School of Integrated Circuits, Peking University
Computer ArchitectureDesign AutomationEmerging Memory
Runsheng Wang
Runsheng Wang
Peking University
M
Meng Li
Institute for Artificial Intelligence, Peking University, Beijing, China; School of Integrated Circuits, Peking University, Beijing, China; Beijing Advanced Innovation Center for Integrated Circuits, Beijing, China