MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high cost and difficulty of reproducing and debugging failures—such as gradient overflow and loss divergence—in reinforcement learning (RL) post-training of large mixture-of-experts (MoE) models. To enable efficient diagnosis, the authors propose a structure-preserving expert pruning method that constructs a lightweight proxy model by clustering and selecting representative experts. This proxy retains the original MoE architecture, routing mechanism, and core capabilities while accurately replicating its training dynamics and failure behaviors. For the first time, this approach enables systematic, low-overhead fault diagnosis in RL training of large MoE models, reducing accelerator requirements by 50%–87.5% and decreasing per-step NPU-hour costs by up to 33.3×, all while maintaining strong consistency in training dynamics.
📝 Abstract
Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead. In practice, factors such as framework adaptation, numerical precision, and operator implementation can cause failures, including gradient overflow and loss divergence. Reproducing such failures directly on large models requires considerable time and computational resources. This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction. Based on these factors, we propose a proxy-model construction method for low-cost fault investigation and auxiliary diagnosis. It employs structure-preserving, clustering-based expert pruning to select representative experts while retaining the model's backbone architecture, routing mechanism, and basic task capabilities. Our experimental results show that the proxy models reduce accelerator requirements by 50%-87.5% and achieve up to a 33.3x reduction in per-step NPU-hour cost, while preserving major training dynamics and reproducing fault responses consistent with the original models. Overall, the proxy models can serve as low-cost surrogates for fault reproduction, targeted validation, and auxiliary diagnosis in RL post-training.
Problem

Research questions and friction points this paper is trying to address.

failure reproduction
RL post-training
large language models
debugging overhead
fault diagnosis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Experts (MoE)
proxy model
failure reproduction
RL post-training
expert pruning
🔎 Similar Papers
No similar papers found.
Y
Yikai Wang
State Key Laboratory for Novel Software Technology, Nanjing University, China
C
Chuansai Zhou
Huawei, China
Y
Yuhang Zhou
State Key Laboratory for Novel Software Technology, Nanjing University, China
W
Weiqiang Wu
Huawei, China
C
Cong Wu
Huawei, China
Y
Yue Deng
State Key Laboratory for Novel Software Technology, Nanjing University, China
B
Ben Feng
Huawei, China
M
Mingming Zhu
Huawei, China
B
Beirong Zhou
Huawei, China
Zhibin Wang
Zhibin Wang
Zhejiang University
new particle formationaerosolshygroscopicityblack carbon
Sheng Zhong
Sheng Zhong
Nanjing University
computer networkssecurity and privacytheory of computing
Chen Tian
Chen Tian
Prof. of Nanjing University
Data Center NetworkingNetwork Function VirtualisationContent Distribution
W
Wangze Zhang
Huawei, China