When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对教师指导可能导致学生模型性能下降的问题,提出了一种基于奖励对齐的在线策略蒸馏方法(RA-OPD),通过筛选与结果奖励一致的轨迹来提高学生模型的表现。
📝 Abstract
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidance on student-generated prefixes is not always reliable. Training should optimize the model to generate responses that are more likely to be correct, or equivalently, to get higher outcome rewards. But during OPD, the teacher model may provide guidance that discourages the student from moving toward correct trajectories or moves the student toward incorrect ones, which is misaligned with outcome reward. Such misaligned guidance is unreliable, as it would mislead the optimization process and ultimately degrade model performance. To mitigate misaligned teacher guidance, we propose Reward-Aligned On-Policy Distillation (RA-OPD). The key insight is to keep only trajectories whose induced updates move the student toward correct trajectories or discourage the student from moving toward incorrect ones. Specifically, for each sampled trajectory, RA-OPD checks whether its trajectory-level distillation return is consistent with its outcome reward and then filters out the misaligned trajectories. RA-OPD selects more reliable trajectories to improve student model performance without requiring additional computational cost. We evaluate RA-OPD on math and code benchmarks using models from the Qwen3 family and the DeepSeek-R1 family. Across seven math benchmarks and three code benchmarks, RA-OPD significantly outperforms standard OPD and other tested OPD variants.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
teacher guidance
misaligned guidance
outcome reward
model performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward-Aligned On-Policy Distillation
trajectory filtering
teacher guidance
outcome reward alignment
model performance improvement
🔎 Similar Papers
No similar papers found.
S
Siyuan Gan
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
Y
Yuhan Li
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China; Shanghai Artificial Intelligence Laboratory, Shanghai, China
X
Xiran Wang
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China; Shanghai Artificial Intelligence Laboratory, Shanghai, China
L
Linjian Meng
Shanghai Artificial Intelligence Laboratory, Shanghai, China
B
Boyan Wang
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
Zhen Zhao
Zhen Zhao
Researcher @ Shanghai AI Lab
Machine LearningAI4ScienceComputer VisionMedical Image Analysis
Jing Huo
Jing Huo
Nanjing University
Machine LearningComputer Vision
Y
Yang Gao
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China