UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决文本引导的无人机地理定位问题,提出UniGeo模型,通过多模态学习和跨视图生成等方法提高定位精度。
📝 Abstract
Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended text query and candidate images. However, incomplete queries and highly similar candidates often make global cross-modal matching insufficient for reliable fine-grained localization. We propose UniGeo, a unified multimodal large language model (MLLM) for text-guided drone geo-localization. Built on a shared vision-language framework, UniGeo jointly supports geo-semantic understanding, cross-view semantic generation, and candidate-level verification. Specifically, it establishes stable correspondences among local scene elements, spatial relations, and language descriptions through geo-semantic learning, and further models semantic mappings between drone and satellite views through cross-view generation. Based on these capabilities, a plug-and-play verification module performs fine-grained discrimination among highly confusable candidates. We further introduce a multi-stage training strategy that progressively learns geo-semantic understanding, cross-view generation, and candidate verification, improving adaptation to text-guided geo-localization. Experiments demonstrate consistent improvements across multiple retrieval backbones. On GeoText-1652, UniGeo improves R@10 and mAP by 13.59 and 2.83 percentage points, respectively, validating its effectiveness for fine-grained text-guided drone geo-localization.
Problem

Research questions and friction points this paper is trying to address.

Text-guided
Drone geo-localization
Cross-modal matching
Fine-grained localization
Semantic understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-modal Large Language Model
Text-guided Geo-localization
Cross-view Generation
Geo-semantic Understanding
🔎 Similar Papers
No similar papers found.