FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入FlowCPO解决了偏好对齐方法间关系不明确的问题,采用离线前向KL目标利用正负样本,无需在线采样。
📝 Abstract
Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on fixed preference pairs rely primarily on positive-only fine-tuning or DPO-style likelihood-ratio surrogates. We organize these approaches through a divergence-based framework and introduce FlowCPO, an offline forward-KL objective that uses both preferred and dispreferred samples without online rollouts. For linear interpolation, we show under explicit regularity conditions that the forward-KL objective is bounded by a contrastive flow matching loss, yielding a tractable surrogate on fixed data. We further show that this loss is nonnegative, whereas the signed regression loss of simplified FlowDPO can be unbounded below. In the in-domain setting, FlowCPO achieves higher mean GenEval and OCR scores than the evaluated baselines, reaching 0.84 and 0.87 versus 0.81 and 0.74 for FlowDPO at CFG 3.0. In the out-of-domain setting, the results are mixed, with the best GenEval result but lower reward scores than RFT on several metrics.
Problem

Research questions and friction points this paper is trying to address.

preference alignment
flow models
offline optimization
forward-process alignment
divergence-based framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

offline forward-KL objective
contrastive flow matching loss
nonnegative loss
🔎 Similar Papers
Y
Yansen Han
Westlake University, Zhejiang University
S
Shengyi Liao
Kling Team, Kuaishou Technology
P
Peng Sun
Westlake University, Zhejiang University
D
Deyuan Liu
Westlake University
Yuanxing Zhang
Yuanxing Zhang
Kuaishou Technology
Recommender SystemLarge Language ModelVideo Understanding
Pengfei Wan
Pengfei Wan
Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
T
Tao Lin
Westlake University