Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于提示级别教师可靠性的验证方法TGOPD,用于改进在线策略蒸馏过程中学生模型的学习效果,并提高了计算资源的利用率。
📝 Abstract
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
teacher reliability
dense supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Teacher Gating
Reliability Verification
Verifier-Scored Teacher Probes
Compute Efficiency
🔎 Similar Papers
No similar papers found.
Z
Zhiwei Zhang
Z
Zechen Sun
F
Fei Zhao
K
Kang Peng
B
Bin Liang
H
Huayu Deng
Yao Hu
Yao Hu
浙江大学
Machine Learning
K
Kam-Fai Wong
M
Mu Chuan