Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of cross-view geolocalization between street-level and satellite imagery, which is hindered by drastic viewpoint shifts and appearance discrepancies. The authors propose a joint view-consensus guided learning framework that operates without explicit geometric warping. Instead of relying on geometric alignment, the method dynamically discovers semantic consistency in feature space through joint-view pathways and employs a global pattern probe to unify the heterogeneous feature metric spaces. Cross-view associations are further enhanced via consensus-guided contrastive learning, and the learned joint-view consensus knowledge is distilled into a single-view encoder to enable efficient inference. Evaluated on four standard benchmarks, the proposed approach significantly outperforms existing methods, achieving more robust and accurate cross-view geolocalization.
πŸ“ Abstract
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.
Problem

Research questions and friction points this paper is trying to address.

cross-view geo-localization
viewpoint changes
appearance discrepancies
geometric warping
semantic consensus
Innovation

Methods, ideas, or system contributions that make the work stand out.

warp-free
feature-space consensus
cross-view geo-localization
joint-view learning
semantic alignment
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.