Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决跨视图地理定位中模型冗余和视觉相似性问题,提出MVLGeo框架,通过多视图专家混合架构、视觉-语言重排序及自适应椭圆先验方法提升定位精度。
📝 Abstract
Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter redundancy and impeding cross-view knowledge sharing. Moreover, top-ranked satellite candidates are often visually similar, so visual appearance and categorical labels alone are insufficient to resolve such ambiguity. To address these, we propose MVLGeo, an efficient framework designed to unify multiple viewpoints and reduce model redundancy. First, we introduce environmental contextual text from the query view as cues to distinguish visually similar candidates via Vision-Language Reranking (VL-Rerank). Second, we design a multi-view Mixture-of-Experts architecture (MV-MoE) with a shared encoder and view-specific experts to reduce redundancy and promote knowledge sharing, while cross-view contrastive learning aligns their representations for consistency. Third, we introduce an adaptive elliptical prior (ESAM-Prior) as auxiliary positional encoding for anisotropic geometric perception. Extensive experiments on the CVOGL benchmarks confirm that MVLGeo, as a unified model for multiple query viewpoints, achieves state-of-the-art performance, demonstrating robustness to input degradation and generalization across viewpoints. Code and models will be available on GitHub to facilitate future work.
Problem

Research questions and friction points this paper is trying to address.

Cross-View Object Geo-Localization
Parameter Redundancy
Knowledge Sharing
Visually Similar Candidates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Reranking
Multi-View Mixture-of-Experts
Adaptive Elliptical Prior
🔎 Similar Papers
2024-06-03International Conference on Machine LearningCitations: 19
X
Xuyu Fan
College of Computer Science, Beijing University of Technology
Q
Qi Ming
College of Computer Science, Beijing University of Technology
Z
Zhu Han
College of Computer Science, Beijing University of Technology
L
Liuqian Wang
Zhengzhou University
S
Siyuan Cao
Zhejiang University
Xiaohan Zhang
Xiaohan Zhang
PhD student of Zhejiang University
Computer VisionObject Detection
X
Xudong Zhao
Beijing Institute of Technology
M
Mingjing Zhao
Beijing Electronic Science and Technology Institute
Yuhan Zhang
Yuhan Zhang
Institute of Automation, Chinese Academy of Sciences