FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出FlowVVTON,通过使用光流作为训练时的监督信号来解决视频虚拟试衣中的大动作和遮挡问题,实现无掩码、多尺度时间一致性。
📝 Abstract
Video virtual try-on aims to transfer a target garment onto a moving person across video frames. Current methods rely on human parsing masks or pose keypoints that frequently fail under large motions and occlusions, causing boundary artifacts and temporal inconsistency. A further limitation is that most approaches rely solely on attention mechanisms for temporal modeling, providing no explicit motion supervision. We propose FlowVVTON, a mask-free framework that eliminates parsing mask dependency entirely. Optical flow is used solely as a training-time supervision signal: a flow-warped latent loss, applied across all layers of the generation model, enforces multi-scale temporal consistency by aligning adjacent-frame features under explicit physical motion constraints. A two-stage training strategy establishes mask-free spatial alignment before introducing flow-guided temporal supervision. Experiments on TikTokDress show that FlowVVTON outperforms baselines by substantial margins, particularly in temporal consistency (5.7$\times$ VFID-R improvement over SwiftTry), while requiring no segmentation masks, pose keypoints, or region annotations at any stage.
Problem

Research questions and friction points this paper is trying to address.

video virtual try-on
large motions
occlusions
temporal inconsistency
motion supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-Guided
Mask-Free
Temporal Consistency
Optical Flow
🔎 Similar Papers
No similar papers found.