SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在混合专家Transformer中循环中间层来提高效率,同时保持计算量匹配,提出SMELT方法,减少了训练FLOPs并提升了性能。
📝 Abstract
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.
Problem

Research questions and friction points this paper is trying to address.

Looped Transformers
Mixture-of-Experts
FLOPs
Scaling Laws
Innovation

Methods, ideas, or system contributions that make the work stand out.

SMELT
Mixture-of-Experts Transformers
Looped Transformers
attention sink
💼 Related Jobs
No related jobs found.
Shaowen Wang
Shaowen Wang
Professor, University of Illinois Urbana-Champaign
CyberGISGeospatial Data ScienceSpatial AISpatial AnalysisSustainability
G
Ge Zhang
ByteDance Seed
K
Kairong Luo
M-A-P
Y
Yuhao Wu
TokenWave.AI
S
Shaofan Liu
Tsinghua University
J
Jiaheng Liu
ByteDance Seed
W
Wenhao Huang
M-A-P
S
Shen Yan
TokenWave.AI
J
Jian Li
Tsinghua University