Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出On-Policy Reverse Distillation方法,旨在通过强化学生模型中验证器驱动的策略梯度部分来解决弱模型指导强模型学习的问题。
📝 Abstract
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.
Problem

Research questions and friction points this paper is trying to address.

Weak-to-Strong Generalization
Model Generations
Multi-Domain Consolidation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Reverse Distillation
Verifier-Driven Policy Gradient
Weak-to-Strong Generalization
🔎 Similar Papers
No similar papers found.