UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing video virtual try-on methods, which rely on multi-stage pipelines that complicate deployment and suffer from irreversible error propagation due to explicit geometric priors. The authors reformulate the task as semantic-conditioned video generation and propose a unified end-to-end architecture that eliminates the need for explicit pose estimation, segmentation, or garment deformation modules. By leveraging a multimodal large language model to implicitly capture dressing semantics and integrating it with a lightweight semantic bridging module within a diffusion-based video generation framework, they devise a three-stage progressive training strategy to effectively couple heterogeneous components. Experiments demonstrate state-of-the-art performance across multiple benchmarks, confirming that implicit semantic guidance can successfully replace fragile geometric preprocessing.
📝 Abstract
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Problem

Research questions and friction points this paper is trying to address.

Video Virtual Try-On
geometric priors
multi-stage pipeline
error propagation
end-to-end generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end framework
semantic conditioning
multimodal large language model
diffusion-based video generation
implicit guidance
💼 Related Jobs
No related jobs found.
Y
Yushe Cao
Tsinghua University
Shikun Feng
Shikun Feng
Baidu
nlp
Fei Shen
Fei Shen
National University of Singapore
Controllable GenerationMultimodal Safety
H
Haikuo Peng
National University of Defense Technology
J
Jianqiang Xia
Shanghai Jiao Tong University
Yiheng Zhu
Yiheng Zhu
Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence
AI for ScienceDeep generative modelsProtein designDrug discovery
D
Dianxi Shi
Tsinghua University
C
Chun Yu
Tsinghua University