Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of current geospatial foundation model evaluations that prioritize accuracy while neglecting calibration and robustness. We propose a comprehensive assessment paradigm transcending single-metric benchmarks by systematically evaluating encoder performance across multiple datasets and distribution shift stress axes using CKA analysis, uncertainty quantification, and selective prediction. Our findings reveal that Earth Observation pretraining induces overconfidence and representation rigidity under distribution shifts, rendering existing calibration methods ineffective and demonstrating that accuracy alone fails to reflect true deployment capability. Consequently, this work identifies critical deficiencies in prevailing benchmarks and advocates for establishing a multidimensional evaluation framework that explicitly incorporates calibration and robustness alongside traditional accuracy metrics to ensure reliable real-world application of geospatial foundation models.
📝 Abstract
Geospatial Foundation Models (GeoFMs) are most commonly ranked and selected by accuracy on standard benchmark conditions via averaged ranks. We show that this protocol is too narrow: the promised deployment in critical EO tasks requires further angles of analysis, mainly calibration, the agreement between a model's confidence and its correctness. Across 16 frozen encoders, four classification and five segmentation datasets, and two orthogonal stress axes, every encoder degrades as corruption intensifies, and the ranking changes as well. Across the four classification benchmarks, EO-pretrained and ImageNet-pretrained encoders are indistinguishable on clean accuracy and clean calibration, and EO pretraining provides no more stability under shift than ImageNet pretraining. Under shift the GeoFMs drift further into overconfidence than the ImageNet-pretrained encoders, at every grade and in every corruption family. A centered kernel alignment (CKA) analysis ties this to representational rigidity: EO-pretrained embeddings move less under corruption while losing just as much task information and remaining overconfident. We apply three commonly explored uncertainty quantification methods and find that temperature scaling and deep ensembles cannot counteract the degradation, while a Gaussian-process probe roughly halves ECE under severe cloud only by tripling it on clean data. In selective prediction experiments, we find that confidence-based abstention cannot defer around confidently wrong predictions, and advocate that benchmark rankings and evaluations should therefore operate across a multitude of conditions and metrics to more holistically evaluate model development progress and close the gap to real world deployment scenarios.
Problem

Research questions and friction points this paper is trying to address.

Geospatial Foundation Models
Calibration
Distribution Shifts
Uncertainty Quantification
Model Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geospatial Foundation Models
Calibration
Distribution Shifts
Uncertainty Quantification
Representational Rigidity
🔎 Similar Papers