Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有方法在优化生成模型时易陷入局部最优的问题,本文提出RA-GRPO框架,通过引入扩散反射和反事实路径合成来改进生成质量。
📝 Abstract
Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based methods often explore inefficiently, making them vulnerable to local optima that may degrade semantic faithfulness and visual realism. To address these challenges, we present Reflection-Aware GRPO (RA-GRPO), a new RL-based preference alignment framework for diffusion generative models. The core idea is to improve "forward" generation by incorporating "backward" reflection during optimization. We first introduce Diffusion Reflection, which rectifies intermediate sampling trajectories by inverting the diffusion process with a weak estimator, guiding latent states toward higher-probability regions of the true data manifold. Furthermore, we introduce Counterfactual Path Synthesis to implicitly distill these rectified trajectories into the policy, enabling the model to internalize the benefits of search-based exploration without incurring inference-time overhead. Extensive experiments on T2I and T2V models demonstrate that RA-GRPO significantly outperforms existing methods, particularly in mitigating reward hacking and improving generalization. The method remains architecture-agnostic and integrates seamlessly with standard pipelines, suggesting a promising direction for stable preference alignment.
Problem

Research questions and friction points this paper is trying to address.

policy gradient
local optima
semantic faithfulness
visual realism
preference alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reflection-Aware GRPO
Diffusion Reflection
Counterfactual Path Synthesis
Preference Alignment
Reinforcement Learning
🔎 Similar Papers
No similar papers found.
J
Junlong Wu
Tsinghua University
J
Jiuzhou Lin
Tsinghua University
Jia Sun
Jia Sun
Hong Kong University of Science and Technology (Guangzhou)
Media arts
B
Boheng Zhang
Kuaishou Technology
H
Huaiqing Wang
Kuaishou Technology
D
Dewen Fan
Kuaishou Technology
H
Houde Liu
Tsinghua University
Q
Qianqian Gan
Kuaishou Technology
F
Fan Yang
Kuaishou Technology
T
Tingting Gao
Kuaishou Technology