Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种计算高效的两步超参数转移框架,通过跨模型宽度和令牌维度转移学习率,解决大规模混合专家模型的优化问题。
📝 Abstract
Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across scaling model widths, and subsequently extrapolating to trillion-token horizons. First, we formulate a Maximal Update Parameterization ($μ$P) adaptation for MoE architectures utilizing Multi-head Latent Attention (MLA) and the Muon optimizer, demonstrating that optimal learning rates transfer consistently across width-scaled models. Second, we extend this transferability along the token dimension by establishing a predictive scaling law. By applying linear regression to the optimal values derived from small proxy models on limited budgets, we successfully extrapolate the ideal learning rate to massive training horizons (e.g., 10 trillion tokens) with high fidelity ($R^2=0.95$). Consequently, this indicates that proxy training on small models is sufficient to determine the optimal learning rate for the extensive training of large-scale MoEs. We apply the proposed methodology to pretrain our foundation model (155B total, 17B active parameters) from scratch, and the stable training and evaluation results validate that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs.
Problem

Research questions and friction points this paper is trying to address.

Mixture-of-Experts
hyperparameter optimization
learning rate
large-scale models
token budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

compute-efficient hyperparameter transfer
Mixture-of-Experts
Maximal Update Parameterization
predictive scaling law
💼 Related Jobs
No related jobs found.
N
Nayeon Kim
Kakao Corp.
H
Hojin Lee
Upstage AI
Y
Yunju Bak
Kakao Corp.
J
Jaesun Park
Kakao Corp.
B
Boseop Kim
Kakao Corp.