The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features

📅 2025-05-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the geospatial representation reliability of GPT-4o and Gemini 2.0 Flash on three critical tasks—geocoding, elevation estimation, and reverse geocoding—using Austrian geographic features. We propose the first zero-shot, multimodal LLM geospatial reasoning evaluation framework, integrating authoritative ground-truth benchmarks (administrative boundaries, DEM-derived elevations, and precise coordinates) for quantitative error analysis, and uniquely distinguishing systematic bias from stochastic error. Results reveal that both models fail to accurately reconstruct federal state boundaries, exhibiting substantial regional inconsistency and factual inaccuracies; Gemini 2.0 Flash demonstrates superior overall accuracy and stability compared to GPT-4o. Crucially, domain-specific geographic knowledge gaps constitute the primary performance bottleneck. The work exposes fundamental limitations in current LLMs’ geospatial cognition and provides empirical grounding for geographic knowledge–enhanced fine-tuning strategies.

Technology Category

Application Category

📝 Abstract
As large language models (LLMs) continue to evolve, questions about their trustworthiness in delivering factual information have become increasingly important. This concern also applies to their ability to accurately represent the geographic world. With recent advancements in this field, it is relevant to consider whether and to what extent LLMs' representations of the geographical world can be trusted. This study evaluates the performance of GPT-4o and Gemini 2.0 Flash in three key geospatial tasks: geocoding, elevation estimation, and reverse geocoding. In the geocoding task, both models exhibited systematic and random errors in estimating the coordinates of St. Anne's Column in Innsbruck, Austria, with GPT-4o showing greater deviations and Gemini 2.0 Flash demonstrating more precision but a significant systematic offset. For elevation estimation, both models tended to underestimate elevations across Austria, though they captured overall topographical trends, and Gemini 2.0 Flash performed better in eastern regions. The reverse geocoding task, which involved identifying Austrian federal states from coordinates, revealed that Gemini 2.0 Flash outperformed GPT-4o in overall accuracy and F1-scores, demonstrating better consistency across regions. Despite these findings, neither model achieved an accurate reconstruction of Austria's federal states, highlighting persistent misclassifications. The study concludes that while LLMs can approximate geographic information, their accuracy and reliability are inconsistent, underscoring the need for fine-tuning with geographical information to enhance their utility in GIScience and Geoinformatics.
Problem

Research questions and friction points this paper is trying to address.

Assessing LLMs' accuracy in geocoding and elevation tasks
Evaluating reliability of LLMs for geographic representations
Identifying systematic errors in LLMs' geospatial data outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates GPT-4o and Gemini 2.0 Flash
Tests geocoding, elevation, reverse geocoding
Highlights need for geographical fine-tuning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Omid Reza Abbasi
Omid Reza Abbasi
University Assistant PostDoc, Paris Lodron University of Salzburg
GeoAIGIScienceSpatial Data Science
F
F. Welscher
Department of Geoinformatics (Z_GIS), Paris-Lodron University Salzburg, Salzburg, Austria
G
Georg Weinberger
Department of Geoinformatics (Z_GIS), Paris-Lodron University Salzburg, Salzburg, Austria
Johannes Scholz
Johannes Scholz
Prof for Geoinformatics, Department of Geoinformatics (Z_GIS), University of Salzburg
Geographic Information ScienceGeoAIspatial OptimizationIndoor GISspatial ABM