Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
📝 Abstract
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Problem

Research questions and friction points this paper is trying to address.

language-model checkpoints
pretraining loss
benchmark scores
downstream training
Innovation

Methods, ideas, or system contributions that make the work stand out.

checkpoint quality
solution density
downstream performance
weight perturbations