SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling

📅 2025-10-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing remote sensing foundation models often rely on specific data sources or training paradigms, limiting their generalizability and flexibility across downstream tasks. To address this, we propose SITS-DECO—the first pure generative decoder-based foundation model tailored for Satellite Image Time Series (SITS). It employs symbolic prompting to unify multi-temporal, multi-modal remote sensing inputs as discrete token sequences, eliminating the need for task- or modality-specific architectural design. Crucially, SITS-DECO discards spatial convolutions and encoder components, focusing exclusively on dense temporal dynamics modeling. Evaluated on the PASTIS-R crop classification benchmark, it achieves superior performance over state-of-the-art remote sensing large models despite significantly fewer parameters. This demonstrates the effectiveness and scalability of the “lightweight + time-centric + symbolic sequence” paradigm for multi-task Earth observation.

Technology Category

Application Category

📝 Abstract
Earth Observation (EO) Foundation Modelling (FM) holds great promise for simplifying and improving the use of EO data for diverse real-world tasks. However, most existing models require additional adaptation before they can be used and are structured rigidly around particular data sources or training approaches. To address this, we take inspiration from large language models, where diverse tasks, both pre-training and downstream, are implicitly captured through next-token prediction over unified token sequences, leveraging the structure and diversity of the training data. We introduce SITS-DECO (Satellite Image Time Series-DECoder Only), a proof-of-concept generative model that applies this unified-sequence framing to EO data. Using a simple GPT-style decoder-only architecture, and demonstrate its ability to perform useful EO tasks (pixel-wise, multi-temporal, multi-modal crop-type classification) in a purely generative framework. Through symbolic prompting, we show that the model can perform multiple supervised and self-supervised tasks within a single unified architecture, without task- or modality-specific adaptation. Despite its simplicity and lack of spatial context, SITS-DECO outperforms much larger EO foundation models on crop-type classification (PASTIS-R) demonstrating that dense temporal sequence modelling is a critical missing ingredient in the current paradigm. This work exemplifies a data-centric modelling paradigm in which capability arises from the diversity and structure of the training data rather than from architectural complexity. SITS-DECO provides a lightweight, practical route to multi-modal, multi-task EO modelling, and a conceptual bridge toward future generative EO foundation models.
Problem

Research questions and friction points this paper is trying to address.

Developing flexible generative models for satellite image time series analysis
Enabling multitask EO modeling without task-specific architectural adaptations
Addressing limitations of rigid EO foundation models through unified token sequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

GPT-style decoder-only architecture for satellite data
Generative framework for multi-task EO modeling
Symbolic prompting enables unified multi-modal tasks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Samuel J. Barrett
LGND AI / Independent Researcher, Canarias, Spain
D
Docko Sow
Tolbi / Independent Researcher, Dakar, Senegal