Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

📅 2025-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Low accuracy and poor generalization of street-view image geolocation in social media and mobile photography scenarios remain challenging. To address this, we propose a fine-tuning-free retrieval-augmented multimodal large language model (RAG-MLLM) framework. Our method employs the SigLIP encoder to construct a large-scale vector database, integrating the EMP-16 and OSV-5M datasets; it dynamically retrieves geographically relevant context to guide an open-weight multimodal LLM for end-to-end geolocation inference. Unlike conventional supervised training paradigms, our framework enables seamless integration of heterogeneous data sources and lightweight deployment. Evaluated on IM2GPS, IM2GPS3k, and YFCC4k benchmarks, it achieves state-of-the-art performance—demonstrating substantial improvements in both absolute localization accuracy and cross-domain generalization. These results empirically validate the effectiveness of retrieval augmentation in modeling geographic semantics for vision-language geolocation.

Technology Category

Application Category

📝 Abstract
Street-level geolocalization from images is crucial for a wide range of essential applications and services, such as navigation, location-based recommendations, and urban planning. With the growing popularity of social media data and cameras embedded in smartphones, applying traditional computer vision techniques to localize images has become increasingly challenging, yet highly valuable. This paper introduces a novel approach that integrates open-weight and publicly accessible multimodal large language models with retrieval-augmented generation. The method constructs a vector database using the SigLIP encoder on two large-scale datasets (EMP-16 and OSV-5M). Query images are augmented with prompts containing both similar and dissimilar geolocation information retrieved from this database before being processed by the multimodal large language models. Our approach has demonstrated state-of-the-art performance, achieving higher accuracy compared against three widely used benchmark datasets (IM2GPS, IM2GPS3k, and YFCC4k). Importantly, our solution eliminates the need for expensive fine-tuning or retraining and scales seamlessly to incorporate new data sources. The effectiveness of retrieval-augmented generation-based multimodal large language models in geolocation estimation demonstrated by this paper suggests an alternative path to the traditional methods which rely on the training models from scratch, opening new possibilities for more accessible and scalable solutions in GeoAI.
Problem

Research questions and friction points this paper is trying to address.

Addresses street-level geolocalization from images using multimodal AI
Overcomes limitations of traditional computer vision for image localization
Eliminates need for expensive model fine-tuning or retraining processes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal large language models integration
Retrieval-augmented generation with geolocation prompts
Vector database construction using SigLIP encoder
🔎 Similar Papers
2024-06-03International Conference on Machine LearningCitations: 19
Y
Yunus Serhat Bicakci
Vocational School of Social Sciences, Marmara University, Istanbul 34865, Turkiye and also with the Geospatial Data Science Group, School of Geographical & Earth Sciences, University of Glasgow, Glasgow G12 8QQ, Scotland, UK
J
Joseph Shingleton
Geospatial Data Science Group, School of Geographical & Earth Sciences, University of Glasgow, Glasgow G12 8QQ, Scotland, UK
Anahid Basiri
Anahid Basiri
Professor in Geospatial Data Science, University of Glasgow
Geospatial Data ScienceMissing DataGeospatial Data QualityGeo-AILocation Based Services