Balancing Emotional Alignment and Semantic Consistency in Image Generation via Reinforcement Learning with Valence-Arousal Anchoring

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过结合连续的情感-唤醒条件、组相对策略优化及中性语义锚点的流匹配图像生成框架,解决了文本到图像生成中情感对齐与语义一致性之间的平衡问题。
📝 Abstract
Continuous emotion control in text-to-image generation requires a model to improve affective alignment without changing the objects, layout, or scene described by the prompt. Existing supervised emotion-injection methods often optimize feature-space proxies and may therefore exhibit emotion-semantic drift, in which stronger emotional conditioning is accompanied by unintended content changes. We address this problem with a flow-matching image-generation framework that combines continuous valence-arousal (VA) conditioning, Group Relative Policy Optimization (GRPO), and a neutral semantic anchor. The deterministic probability-flow ODE is converted into a marginal-preserving SDE, yielding non-degenerate transition densities for trajectory sampling and policy-ratio estimation. A frozen CLIP-based VA regressor supplies a terminal reward measuring the distance between the predicted and target VA coordinates, while an image generated from the same prompt under zero VA conditioning provides a feature-space reference for semantic preservation. A reduced denoising schedule is used for online RL sampling, whereas the original schedule is retained at inference. Experiments on 3,300 prompt-emotion combinations show substantially lower valence and arousal errors than the VA-conditioned baseline and an improved CLIPScore relative to EmotiCrafter, with a measurable trade-off in reference-free image quality. The results support anchor-regularized Flow-GRPO as a practical approach to balancing emotional alignment and semantic consistency in continuous-affect image synthesis.
Problem

Research questions and friction points this paper is trying to address.

emotional alignment
semantic consistency
text-to-image generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Valence-Arousal Conditioning
Group Relative Policy Optimization (GRPO)
Flow-Matching Framework
Semantic Anchoring
Continuous Emotion Control
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jisheng Dang
School of Information Science & Engineering, Lanzhou University, Lanzhou 730000, China
Z
Zhenxuan Wang
School of Information Science & Engineering, Lanzhou University, Lanzhou 730000, China
B
Bin Li
School of Environmental and Spatial Informatics, China University of Mining and Technology, Xuzhou 221116, China
Ronghao Lin
Ronghao Lin
University of Science and Technology of China
Waveform DesignSparse Array DesignStatistical Signal ProcessingOptimization Theory.
B
Bin Hu
School of Medical Technology, Beijing Institute of Technology, Beijing 100081, China
T
Tat-Seng Chua
School of Computing, National University of Singapore, Singapore 119077