Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决一步生成模型的奖励引导微调问题,本文通过Wasserstein梯度流提出了一种新的微调方法,适用于不同类型的奖励,并在多个数据集上验证了其有效性。
📝 Abstract
To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. To address this, we consider one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space. We then propose a novel reward-guided fine-tuning of a one-step generative model via WGF. We derive a practical training method that requires no reward gradients, thereby handling both non-differentiable and differentiable rewards. Moreover, our method provides smooth and stable reward-guided distributional updates while mitigating reward hacking and mode collapse. Experiments on 2D synthetic data, CIFAR-10, and ImageNet 256$\times$256 with diverse rewards, including JPEG (in)compressibility, class probability, Black-and-White and CLIP alignment, show that our method achieves better reward alignment compared to baselines.
Problem

Research questions and friction points this paper is trying to address.

one-step generative models
reward-guided fine-tuning
Wasserstein Gradient Flow
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wasserstein Gradient Flow
reward-guided fine-tuning
one-step generative models
distributional evolution
🔎 Similar Papers
No similar papers found.