CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对跨家族在线策略蒸馏效果不佳的问题,提出CompassOPD方法,通过消除偏差并传递同家族内似然偏移来提高学生模型性能。
📝 Abstract
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger external teachers offering little additional improvement. To understand this disconnect, we decompose the cross-family OPD signal into two components: an offset between a low-capability teacher-family reference and the student, and the within-family log-likelihood shift from that reference to the strong teacher. Standard OPD transfers both components together, allowing the offset to dominate the update direction and obscure the changes associated with teacher capability improvements. We propose CompassOPD, which removes this offset and transfers the within-family shift, while a frozen student reference anchors updates to the student's initial policy. Thus, both teacher-side and student-side changes are measured within their respective model families. Experiments across three student families and multiple teacher families show that CompassOPD consistently outperforms standard cross-family OPD, improving average reasoning accuracy by up to 5.50 points. For an MoE teacher, we further construct the reference directly from the teacher checkpoint by reducing expert activation, eliminating the need for a separate reference checkpoint while retaining a 3.43-point gain over OPD.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Cross-family settings
Tokenizer alignment
Model family
Student-generated trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Family On-Policy Distillation
Within-Family Likelihood Shifts
Frozen Student Reference
Model Family Alignment
🔎 Similar Papers
No similar papers found.
N
Naibin Gu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Q
Qingyi Si
JD.COM
Chenxu Yang
Chenxu Yang
Institute of Information Engineering, Chinese Academy of Sciences
NLPDialogue Generation
C
Chuanyu Qin
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
J
Junhao Zhou
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Peng Fu
Peng Fu
Institute of Information Engineering, Chinese Academy of Sciences
Natural Language Processing
Zheng Lin
Zheng Lin
Institute of Information Engineering, CAS
NLP
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security