Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频生成模型与人类偏好对齐的高计算成本问题,提出了一种基于分布匹配的统一单阶段优化框架DM-Align,通过结合两个梯度方向来提升生成质量和偏好一致性。
📝 Abstract
Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applying RL after distillation frequently leads to model collapse. To overcome these limitations, we propose a unified, single-stage optimization framework grounded in Distribution Matching (DM). In the standard DM framework, distillation updates the model via a gradient direction that minimizes the gap between the real and fake models, guiding generations toward clarity and high fidelity. Building upon this, we introduce DM-Align, which derives a complementary gradient direction to guide the model toward human-preferred samples. Inspired by DPO and GRPO, our method leverages the distributional gap -- formulated from either preference pairs or intra-group exploration -- to directly construct this preference-guided gradient. By synergizing these two gradient directions, our approach eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL. Comprehensive experiments across multiple foundational video models demonstrate that this sample-guided framework robustly enhances both distillation quality and preference alignment, consistently outperforming both standalone variants and sequential two-stage pipelines.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Distribution Matching
Video Generation
Model Distillation
Human Preferences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distribution Matching
DM-Align
Preference-guided Gradient
🔎 Similar Papers
J
Jiuzhou Lin
Tsinghua University
J
Junlong Wu
Tsinghua University
Fei Zuo
Fei Zuo
University of Central Oklahoma
System and Network SecurityMachine LearningInternet of Things
H
Huan Ouyang
Beijing University of Posts and Telecommunications
D
Dewen Fan
Kuaishou Technology
B
Boheng Zhang
Kuaishou Technology
H
Huaiqing Wang
Kuaishou Technology
Jia Sun
Jia Sun
Hong Kong University of Science and Technology (Guangzhou)
Media arts
F
Fan Yang
Kuaishou Technology
H
Houde Liu
Tsinghua University
Kehai Chen
Kehai Chen
Harbin Institute of Technolgy (Shenzhen)
LLMNatural Language ProcessingAgentMulti-model Generation
M
Min Zhang
Harbin Institute of Technology (ShenZhen)
T
Tingting Gao
Kuaishou Technology
H
Han Li
Kuaishou Technology