TryOnReward: Learning Foveated Consistency for Reinforcement Fine-Tuning of Virtual Try-On

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决虚拟试穿中评分函数与人类偏好不一致的问题,提出TryOnReward模型,通过细粒度奖励机制和视觉-语言基础优化试穿效果。
📝 Abstract
Virtual Try-On (VTON) aims to dress a person with the reference garment, producing visually reasonable results aligned with human preferences. Turning this preference-oriented goal into an actionable objective relies on a scoring function aligned with human taste. However, classic fidelity metrics exhibit weak correlation with human judgments, and generic VLMs fail to provide the discriminative granularity demanded by try-on quality evaluation, which hinges on faithfully preserving garment and person details. This shortcoming is further exacerbated in the reinforcement fine-tuning (RFT) optimization and leads to severe reward hacking. To this end, we present TryOnReward, a fine-grained reward model tailored for VTON. Built on a vision-language backbone, it adopts a foveation calibration objective that grounds each quality dimension in the relevant region to avoid global shortcut learning. Meanwhile, TryOnReward jointly optimizes pairwise preferences and per-dimension quality scores via margin-aware supervision, leveraging both relative and absolute quality signals. For model training and evaluation, we build TryOnReward-100K, a human-annotated per-dimension rating dataset, alongside TryOn-Bench and TryOnRewardBench, two benchmarks covering diverse real scenarios. Extensive experiments confirm that TryOnReward significantly outperforms generic judges in human preference alignment, and when serving as the RFT reward function, it consistently yields human-preferred try-on results across multiple baselines.
Problem

Research questions and friction points this paper is trying to address.

Virtual Try-On
Reinforcement Fine-Tuning
Fidelity Metrics
Vision-Language Models
Reward Hacking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foveation Calibration
Reinforcement Fine-Tuning (RFT)
Pairwise Preferences
Margin-Aware Supervision
Vision-Language Model
X
Xueheng Li
JD.com, China; University of Science and Technology of China, China
Y
Yong Liu
JD.com, China
X
Xiaolong Fu
JD.com, China
W
Wen Xue
JD.com, China; South China University of Technology, China
C
Chengjun Xie
University of Science and Technology of China, China
Yipeng Sun
Yipeng Sun
Friedrich-Alexander-Universität Erlangen-Nürnberg
Deep LearningImage ProcessingInverse Problem
Y
Yan Li
JD.com, China
S
Simiu Gu
JD.com, China