PDA++: Field-Aligned Planning and Scene-Adaptive Insertion in Remote Sensing
为解决遥感识别中稀有目标观察不足和标注成本高的问题,提出PDA++框架,通过环境感知的对象插入方法提高目标多样性并适应背景环境。
为解决遥感识别中稀有目标观察不足和标注成本高的问题,提出PDA++框架,通过环境感知的对象插入方法提高目标多样性并适应背景环境。
为解决遥感模型改进依赖手动的问题,提出RingMoClaw框架,通过多代理和经验循环优化模型性能,减少迭代步骤。
Remote sensing image fusion models suffer from limited generalizability due to scarce authentic ground-truth data and domain shifts across heterogeneous sensors. To address this, we introduce, for the first time, a foundation model paradigm into remote sensing fusion, proposing a large-scale pretraining framework grounded in spatial-spectral priors. Our method synthesizes a highly diverse dataset by applying realistic degradations—including blur, noise, and downsampling—to ImageNet and SkyScript images. This enables effective pretraining across multiple architectures, including CNNs, Transformers, and Mamba. The resulting model achieves zero-shot and few-shot pan-sharpening, outperforming state-of-the-art methods on six major satellite datasets (e.g., WorldView). Remarkably, it adapts to unseen sensor domains with fine-tuning on merely a single real-world image. Our work establishes a new benchmark for cross-domain generalization in remote sensing image fusion.
Multi-temporal remote sensing change detection suffers from fine-grained recognition difficulties caused by heterogeneity and spatiotemporal misalignment, while existing sequential modeling approaches often compromise local structural consistency. To address this, we propose a structure-aware interleaved state-space modeling framework. First, we introduce a novel checkerboard serpentine scanning strategy that preserves local structural integrity across multi-temporal features and enables single-pass forward alignment. Second, a multi-dilation convolution fusion module explicitly captures center-to-corner contextual relationships, enhancing robustness to misalignment. Third, we integrate a SpatialMamba encoder with a lightweight cross-source interaction module for efficient heterogeneous temporal feature fusion. Our method achieves state-of-the-art performance on binary change detection, semantic change detection, and multimodal building damage assessment—demonstrating significant improvements in change localization accuracy and cross-scenario generalization.
Monocular depth estimation (MDE) is inherently ill-posed, making reliable 3D structure recovery from a single 2D image challenging. To address this, we propose WEDepth—a zero-shot adaptation method that activates geometric and scene priors embedded in pre-trained vision foundation models (VFMs) without fine-tuning, architectural modification, or weight updates. WEDepth employs a multi-level feature injection mechanism that explicitly integrates shallow texture and deep semantic features, enabling context-aware depth reconstruction. Evaluated on NYU-Depth v2 and KITTI, it achieves state-of-the-art performance—comparable to diffusion-based or multi-step inference methods—while demonstrating strong cross-domain zero-shot generalization. Our key contribution is the first plug-and-play activation of multi-level implicit priors within VFMs, achieving an unprecedented balance among computational efficiency, architectural generality, and interpretability.
为解决遥感识别中稀有目标观察不足和标注成本高的问题,提出PDA++框架,通过环境感知的对象插入方法提高目标多样性并适应背景环境。
为解决遥感模型改进依赖手动的问题,提出RingMoClaw框架,通过多代理和经验循环优化模型性能,减少迭代步骤。
Remote sensing image fusion models suffer from limited generalizability due to scarce authentic ground-truth data and domain shifts across heterogeneous sensors. To address this, we introduce, for the first time, a foundation model paradigm into remote sensing fusion, proposing a large-scale pretraining framework grounded in spatial-spectral priors. Our method synthesizes a highly diverse dataset by applying realistic degradations—including blur, noise, and downsampling—to ImageNet and SkyScript images. This enables effective pretraining across multiple architectures, including CNNs, Transformers, and Mamba. The resulting model achieves zero-shot and few-shot pan-sharpening, outperforming state-of-the-art methods on six major satellite datasets (e.g., WorldView). Remarkably, it adapts to unseen sensor domains with fine-tuning on merely a single real-world image. Our work establishes a new benchmark for cross-domain generalization in remote sensing image fusion.
Multi-temporal remote sensing change detection suffers from fine-grained recognition difficulties caused by heterogeneity and spatiotemporal misalignment, while existing sequential modeling approaches often compromise local structural consistency. To address this, we propose a structure-aware interleaved state-space modeling framework. First, we introduce a novel checkerboard serpentine scanning strategy that preserves local structural integrity across multi-temporal features and enables single-pass forward alignment. Second, a multi-dilation convolution fusion module explicitly captures center-to-corner contextual relationships, enhancing robustness to misalignment. Third, we integrate a SpatialMamba encoder with a lightweight cross-source interaction module for efficient heterogeneous temporal feature fusion. Our method achieves state-of-the-art performance on binary change detection, semantic change detection, and multimodal building damage assessment—demonstrating significant improvements in change localization accuracy and cross-scenario generalization.
Monocular depth estimation (MDE) is inherently ill-posed, making reliable 3D structure recovery from a single 2D image challenging. To address this, we propose WEDepth—a zero-shot adaptation method that activates geometric and scene priors embedded in pre-trained vision foundation models (VFMs) without fine-tuning, architectural modification, or weight updates. WEDepth employs a multi-level feature injection mechanism that explicitly integrates shallow texture and deep semantic features, enabling context-aware depth reconstruction. Evaluated on NYU-Depth v2 and KITTI, it achieves state-of-the-art performance—comparable to diffusion-based or multi-step inference methods—while demonstrating strong cross-domain zero-shot generalization. Our key contribution is the first plug-and-play activation of multi-level implicit priors within VFMs, achieving an unprecedented balance among computational efficiency, architectural generality, and interpretability.