Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对流偏好优化中的流形漂移问题,提出了一种温度控制的方法ThermoDPO及其加权变体,以保持生成模型的样本在预训练数据流形上。
📝 Abstract
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify transport trajectories without an inherent constraint to the pretrained data manifold and can move terminal samples off the pretrained support. We formalize this failure mode as manifold drift. Theoretically, we show that optimal flow matching recovers the terminal data distribution, whereas a preference update leaves the pretrained manifold whenever its induced terminal displacement has a nonzero normal component. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO and controls a pointwise reconstruction-based surrogate for manifold distance. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted. On the main toy benchmark, ThermoDPO-weighted attains a StrictScore of 0.899, compared with 0.629 for FlowDPO and 0.857 for FlowDPO+RFT. On SD3.5-M at CFG = 4.5, it improves OCR by 47.5% and the average of four metrics by 16.0%.
Problem

Research questions and friction points this paper is trying to address.

manifold drift
flow preference optimization
reward-driven updates
pretrained data manifold
terminal samples
Innovation

Methods, ideas, or system contributions that make the work stand out.

ThermoDPO
Manifold Drift
Flow Matching
Temperature-controlled Objective
💼 Related Jobs
No related jobs found.