Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Self-OPD,一种无教师的在线策略蒸馏框架,通过自探索生成监督信号,优化流匹配模型,解决高计算成本和分布差异问题。
📝 Abstract
On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Models (LLMs) and has recently been adapted to flow matching models. However, this paradigm suffers from two major issues: First, training a separate, task-specific teacher for every new objective incurs high computational costs. Second, the discrepancy between teacher and student distributions often leads to compounding errors along the generation trajectory. In this paper, we introduce \textbf{Self-OPD}, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision. At each timestep, Self-OPD branches the deterministic next-state prediction into $K$ stochastic SDE candidates, rolls them out with the ODE sampler, and compares their rewards against a deterministic self-reference baseline to obtain normalized advantages. The velocity field is optimized with an all-branch pull-push objective, where high-advantage branches attract the student and low-advantage branches repel it under direction-aware attenuation and SDE-variance normalization. For multi-objective alignment, Self-OPD fuses normalized scores at the reward level, avoiding direct gradient conflict. Experiments on single and mixed reward benchmarks show that Self-OPD outperforms prior RL and OPD methods without task-specific teachers.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Flow Matching Models
Computational Costs
Distribution Discrepancy
Compounding Errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-OPD
on-policy distillation
flow matching models
stochastic SDE candidates
all-branch pull-push objective
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Shiyi Zhang
Shiyi Zhang
Tsinghua University
Video GenerationVideo Understanding
Mushui Liu
Mushui Liu
Zhejiang University
Generative ModelsMulti-modal LearningFew-shot Learning
Y
Yunze Tong
Zhejiang University
Wanggui He
Wanggui He
Researcher, Alibaba Group
ai
S
Siyu Zou
Alibaba Group
J
Jinlong Liu
Alibaba Group
Y
Yunlong Yu
Zhejiang University
J
Jian Song
Tsinghua University
Hao Jiang
Hao Jiang
Alibaba Group
LLM & AIGC
P
Pipei Huang
Alibaba Group
B
Bo Zheng
Alibaba Group