WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the training instability in conventional online policy distillation caused by joint updates of policy and states. To mitigate this issue, the authors propose a hybrid-constrained co-training framework featuring two trainable policies: an anchor policy that generates trajectories and an auxiliary policy that evaluates states. The outputs of these policies are geometrically mixed to align with a frozen teacher model, and optimization is performed using reverse KL divergence. This approach provides branch-level flexibility unattainable with static targets. Experiments on Qwen3-1.7B and Qwen3-4B demonstrate substantial performance gains, achieving MATH500 accuracies of 0.585 and 0.685, respectively. Moreover, the method effectively curbs entropy growth and trajectory degradation in code generation tasks, yielding state-of-the-art development set scores.
📝 Abstract
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch of offline distillation. The same feedback loop can nevertheless be unstable: each update changes both the policy and the states on which the next update is computed. We introduce WDL-OPD, a mixture-constrained co-training method with two trainable policies. An anchor policy generates every rollout, an auxiliary policy evaluates the same visited states, and a geometric mixture of their token distributions is matched to a frozen teacher by reverse KL. Both policies receive gradient. We show that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express. In recorded Qwen3 experiments at 1.7B and 4B scale, WDL-OPD produces the strongest student checkpoint in each of four scale-domain settings. It raises MATH500 accuracy from 0.630 to 0.685 at 4B and from 0.521 to 0.585 at 1.7B. In code generation, seven single-policy OPD configurations exhibit entropy growth or trajectory degradation, while co-training reaches independently re-evaluated development scores of 0.637 and 0.375. Because several comparisons differ in curriculum or initialization, these results support a stabilization hypothesis rather than a universal causal claim. We provide the exact training algorithm, failure evidence, and the controlled comparison matrix needed to test that hypothesis.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
training instability
policy learning
state distribution shift
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Mixture-constrained co-training
Reverse KL divergence
Anchor-auxiliary policy
Trajectory stabilization
🔎 Similar Papers