Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited understanding of the capability boundaries, cross-task generalization, and task-specific limitations of remote sensing multimodal large language models (RS-MLLMs). It presents the first systematic evaluation comparing general-purpose computer vision multimodal large language models (CV-MLLMs) with domain-specific RS-MLLMs across diverse remote sensing scene understanding tasks, integrating model architecture analysis, multimodal mechanism comparison, and comprehensive benchmarking on varied remote sensing datasets. The findings reveal that generic CV-MLLMs—without any remote sensing fine-tuning—often match or even surpass specialized models, demonstrating strong transferability. Moreover, the work uncovers common bottlenecks across current models in spatial reasoning, fine-grained comprehension, high-resolution image processing, and handling diverse instruction formats.
📝 Abstract
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing Image Understanding
Multimodal Large Language Models
Domain-Specific Models
General-Purpose Models
Cross-Task Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Remote Sensing Image Understanding
Domain-Specific vs General-Purpose
Cross-Task Generalization
Visual Question Answering
🔎 Similar Papers
No similar papers found.
Q
Qiwei Ma
School of Artificial Intelligence and Robotics, Hunan University, Changsha, 410082, China
C
Chunping Qiu
Intelligent Game and Decision Lab (IGDL), Beijing, 100091, China
X
Xinjun Cheng
Intelligent Game and Decision Lab (IGDL), Beijing, 100091, China
X
Xiaoyu Zhang
Intelligent Game and Decision Lab (IGDL), Beijing, 100091, China
P
Puhong Duan
School of Artificial Intelligence and Robotics, Hunan University, Changsha, 410082, China
Ke Yang
Ke Yang
Beijing University of Technology
statisticsmeta-analysis
X
Xudong Kang
School of Artificial Intelligence and Robotics, Hunan University, Changsha, 410082, China, Yuelushan Center for Industrial Innovation, Changsha, 410082, China
S
Shutao Li
School of Artificial Intelligence and Robotics, Hunan University, Changsha, 410082, China