JewelTry: Mask-Free Scale Aware Jewelry Virtual Try-On

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决珠宝虚拟试戴中的尺寸和位置准确性问题,本文提出了一种无需遮罩的扩散框架JewelTry,并引入了JVTO-Bench基准数据集以提供规模忠实度支持。
📝 Abstract
Virtual try-on (VTON) enables customers to visualize how fashion products appear when worn and has become an important technology for online shopping. While recent advances have substantially improved garment VTON, jewelry remains a challenging and underexplored category due to its small size, rigid structure, and sensitivity to fine-grained visual details. Realistic jewelry VTON requires not only faithful appearance transfer but also accurate scale and placement relative to the wearer. Existing jewelry VTON methods typically rely on mask guidance, whereas mask-free approaches lack explicit guidance for modeling the product scale. To bridge this gap, we introduce JVTO-Bench, a benchmark dataset for scale-faithful jewelry VTON, providing reference source target triplets with real-world product-scale annotations across four major jewelry categories. Building upon this benchmark, we propose JewelTry, a mask-free diffusion framework for scale-aware jewelry VTON. JewelTry incorporates a scale adapter that encodes product dimensions into a scale token, enabling the model to learn scale relationships between jewelry items and surrounding human anatomy in-context. To further improve jewelry consistency, we introduce a single-directional condition attention mechanism and an attention refinement loss that preserve both coarse geometry and fine-grained structural details of the reference jewelry. Extensive experiments show that JewelTry achieves a balance among visual fidelity, background preservation, object consistency and scale accuracy, establishing a strong baseline for mask-free, scale-aware jewelry virtual try-on.
Problem

Research questions and friction points this paper is trying to address.

Virtual Try-On
Jewelry
Scale-Aware
Mask-Free
Product Scale
Innovation

Methods, ideas, or system contributions that make the work stand out.

mask-free diffusion framework
scale adapter
single-directional condition attention mechanism
attention refinement loss
💼 Related Jobs
No related jobs found.
X
Xinlei Niu
Australian National University, Canberra, Australia
Peixia Li
Peixia Li
The University of Sydney
computer vision and deep learning
J
Jun Wang
Amazon, Melbourne, Australia
Chenchen Xu
Chenchen Xu
Amazon
Machine LearningNatural Language ProcessingVision Language
Jiayu Yang
Jiayu Yang
The Australian National University
3D Computer Vision3D AIGC3D ReconstructionMulti-view StereoVR AR XR
J
Jing Zhang
Australian National University, Canberra, Australia
Pulak Purkait
Pulak Purkait
Amazon
Computer VisionMachine LearningImage Processing
H
Hongdong Li
Australian National University, Canberra, Australia; Amazon, Melbourne, Australia