DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出DualOPSD,通过自适应调整教师和学生模型来改进在线自蒸馏方法,提高模型在数学竞赛数据上的表现。
📝 Abstract
On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and output style change during training. We propose DualOPSD, an asymmetric alternating framework that adapts both policies. The student first learns from the privileged teacher. The teacher then moves toward the updated student distribution on the same student trajectory. This update makes later supervision responsive to the learner and does not require another rollout. On Qwen3-8B in non-thinking mode, DualOPSD improves avg@12 over OPSD by 23.61, 13.89, and 10.00 points on AIME 2024, AIME 2025, and HMMT 2025. Results at 1.7B and 4B show that the accuracy gain depends on model scale. Across all three scales, DualOPSD reduces truncation. The 4B diagnostic also shows lower KL in both directions between the teacher and student.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
privileged teacher
student distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Privileged Teachers
Asymmetric Alternating Framework
On-Policy Self-Distillation
Student-Teacher Adaptation
Non-Additional Rollout
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1