🤖 AI Summary
This study addresses the limitations of existing remote sensing vision-language models in handling long-temporal Earth observation tasks, such as multi-stage geographic evolution modeling, spatial change localization, and temporal anomaly detection. To this end, the authors introduce LongEarth-Bench, the first benchmark specifically designed for long-sequence remote sensing reasoning, comprising approximately 120,000 question-answer pairs with an average sequence length of 15.14 frames. They further propose LongEarth-R1, a novel model enhanced through supervised fine-tuning, explicit sequence identifiers, and structured chain-of-thought training, augmented by a grouped relative policy optimization framework incorporating spatiotemporal format rewards. This approach substantially improves long-sequence reasoning capabilities, enabling LongEarth-R1 to achieve state-of-the-art performance across all 12 long-temporal tasks while maintaining competitive results on standard remote sensing benchmarks.
📝 Abstract
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.