🤖 AI Summary
This study addresses the limited understanding of the capability boundaries, cross-task generalization, and task-specific limitations of remote sensing multimodal large language models (RS-MLLMs). It presents the first systematic evaluation comparing general-purpose computer vision multimodal large language models (CV-MLLMs) with domain-specific RS-MLLMs across diverse remote sensing scene understanding tasks, integrating model architecture analysis, multimodal mechanism comparison, and comprehensive benchmarking on varied remote sensing datasets. The findings reveal that generic CV-MLLMs—without any remote sensing fine-tuning—often match or even surpass specialized models, demonstrating strong transferability. Moreover, the work uncovers common bottlenecks across current models in spatial reasoning, fine-grained comprehension, high-resolution image processing, and handling diverse instruction formats.
📝 Abstract
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.