Unsupervised Post-Training of Foundation Models: A Survey

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究无监督后训练方法,利用模型自身生成的学习信号而非外部数据,在无标签输入上进行适应性更新,以避免错误放大。
📝 Abstract
Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. We catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator. Beyond inventory, we show how the choice of internal signal and task structure determines whether post-training improves the model or recursively amplifies error. An orthogonal Input Visibility $\times$ Update Persistence view maps deployment regimes and defines a unified framework for UPT selection and evaluation.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Post-Training
Foundation Models
Unlabeled Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Post-Training
Internal Signal
Model Artifacts
Update Persistence
Input Visibility
🔎 Similar Papers