XDG: Accelerated Visual Disambiguation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉混淆问题,XDG通过轻量级LoRA适配器微调Depth Anything 3,并利用紧凑的MLP头部进行图像对分类,实现高效且准确的大规模SfM重建。
📝 Abstract
Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically distinct surfaces can produce incorrect image matches and degrade reconstruction quality. Previous work mitigates this issue with geometry-aware foundation-model features, but places a heavy transformer classifier on top of the backbone, making large-scale disambiguation expensive. We introduce XDG, an efficient visual disambiguation model designed for scalable SfM. Our key observation is that a 3D foundation model already performs the cross-view geometric reasoning necessary for visual disambiguation, so doppelganger classification should adapt the backbone representation directly rather than relearn pair reasoning in a separate heavy decoder. XDG fine-tunes Depth Anything 3 with lightweight LoRA adapters and repurposes its camera tokens as compact pair-level classification tokens. A compact MLP head predicts whether a candidate image pair observes the same 3D surface. Extensive experiments show that XDG provides a favorable accuracy-efficiency tradeoff: it remains competitive with the state-of-the-art disambiguation method across pairwise and reconstruction benchmarks and delivers more than a 3x inference speedup. On individual LaMAR scenes containing thousands of images, XDG saves more than 10 hours of visual disambiguation processing. Code is available at https://github.com/xtcpete/xdg.
Problem

Research questions and friction points this paper is trying to address.

visual aliasing
doppelganger problem
structure-from-motion
image matches
reconstruction quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual disambiguation
lightweight LoRA adapters
camera tokens
compact MLP head
inference speedup
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Gonglin Chen
Gonglin Chen
University of Southern California
Computer VisionMachine Learning
Ben Southall
Ben Southall
Senior Computer Scientist, SRI International
Computer Vision
H
Hanyuan Xiao
USC Institute for Creative Technologies, University of Southern California
Wenbin Teng
Wenbin Teng
University of Southern California
Computer VisionGenerative Model3D reconstruction
H
Haolin Xiong
USC Institute for Creative Technologies, University of Southern California
T
Tianwen Fu
USC Institute for Creative Technologies, University of Southern California
J
Junyi Ouyang
USC Institute for Creative Technologies, University of Southern California
K
Kshitij Singh Minhas
SRI International
Supun Samarasekera
Supun Samarasekera
Technical Director - Vision and Robotics Lab, SRI International
Computer Vision3D ModelingAugmented RealityNavigation
R
Rakesh Kumar
SRI International
Yajie Zhao
Yajie Zhao
Computer Scientist at University of Southern California
Virtual HumanNeural RenderAR/VR