GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入GeoAgent,利用具身导航改进视觉语言模型在地理位置定位中的表现,解决了静态图像分析的局限性。
📝 Abstract
Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied navigation, where a multimodal agent autonomously explores its surroundings to gather observations before submitting a prediction. To this end, we introduce \textbf{GeoAgent}, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning. Our analysis shows that modern VLMs struggle to discern regional patterns while succeeding at country- and continent-level predictions. When compared to static image-based baselines, agentic navigation significantly improves accuracy across established metrics. We also note severe bias in a developed/developing region context across frontier model architectures and poor self-improvement capabilities given incorrect priors. Overall, our work establishes the challenges of embodied navigation and geospatial reasoning. We publicly release our code and the GeoAgent environment: https://geoagent-benchmark.github.io
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
geolocalization
embodied navigation
regional patterns
bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

embodied navigation
multimodal agent
geolocalization
sequential reasoning
Street View
🔎 Similar Papers
No similar papers found.