VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出VISTA方法,通过验证过的rollout调整教师模型以适应学生分布,解决标准OPSD中单向监督可能误导学生的问题。
📝 Abstract
On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, standard OPSD treats the teacher distribution as a fixed target along the student's rollout and updates only the student %, although -- even though privileged conditioning does not guarantee that the teacher always provides the most appropriate target for problem-only reasoning. This one-way supervision can therefore misdirect the student when the teacher distribution is misaligned with valid student reasoning. We therefore introduce Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution. Within each verified rollout, VISTA further restricts this adaptation to the top-$k$ positions with the largest teacher--student KL divergence. Notably, VISTA reuses the rollout and loss function from standard OPSD, introducing no additional sampling or separate reward objective. Across AIME24, AIME25, and HMMT25 with Qwen3 models at 1.7B, 4B, and 8B, VISTA achieves the highest Avg@12 at every scale, improving over OPSD by $0.6$, $0.7$, and $2.1$ points, respectively. These results demonstrate the value of student supervision from outcome-verified rollouts and highlight student-to-teacher adaptation as a promising direction for OPSD.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
student-teacher adaptation
privileged conditioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifier-Informed Adaptation
KL Divergence
Outcome-Verified Rollouts
🔎 Similar Papers
No similar papers found.
Z
Zewen Ding
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
Z
Zezhong Wu
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
Z
Zhou Tao
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
Shida Wang
Shida Wang
National University of Singapore
Sequence ModellingLarge Language Model
S
Shizhuo Hou
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
Y
YongXiang Hua
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
H
Haoyu Cao
University of Science and Technology of China, State Key Laboratory of Cognitive Intelligence
Linli Xu
Linli Xu
University of Science and Technology of China