TV-Regulated OPD: Direction Matters in On-Policy Distillation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对在策略蒸馏中的高方差和噪声问题,提出了一种基于总变差(TV)调控的优势函数平滑方法TV-OPD,以稳定训练过程并保持性能。
📝 Abstract
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the Total Variation (TV) and propose a robust TV regulated On-Policy Distillation (TV-OPD) method. Benefiting from the bounded and diminished advantages, TV-OPD exhibits stable training dynamics and steady late-stage performance. We conducted comprehensive experiments and found that, across various settings, TV-OPD consistently achieved better performance and lower variance in the late-stage of training.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Large Language Models
high variance
noise
training instability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Total Variation
On-Policy Distillation
Stability
Advantages Shaping
🔎 Similar Papers
No similar papers found.
Han Xiao
Han Xiao
Eindhoven University of Technology
Human-Computer Interaction
Yifan Niu
Yifan Niu
PhD student, Hong Kong University of Science and Technology
Machine Learning
D
Dongyi Liu
The Hong Kong University of Science and Technology (Guangzhou)
C
Chang Luo
The University of Edinburgh
J
Jia Li
The Hong Kong University of Science and Technology