π€ AI Summary
This study addresses the unsupervised fusion of low-resolution hyperspectral images with panchromatic images, aiming to simultaneously preserve spatial details and spectral fidelity. To this end, the authors propose a Dual-modality Prompt Diffusion Model (DIDM), which, for the first time, encodes panchromatic and hyperspectral observations as spatial and spectral prompt tokens, respectively, and injects them into intermediate layers of a frozen pre-trained remote sensing diffusion model via cross-attention mechanisms to guide the generation process. The method innovatively incorporates a panchromatic-guided weighted pixel-aware total variation regularizer to balance structural preservation and noise suppression. Experiments on the Pavia, Chikusei, and Houston datasets demonstrate that the proposed approach achieves state-of-the-art performance in terms of the HQNR metric under full-resolution (FR1) evaluation, significantly outperforming existing methods.
π Abstract
Hyperspectral pansharpening aims to reconstruct a high resolution hyperspectral (HRHS) image from a panchromatic (PAN) image and a low resolution hyperspectral (LRHS) image while preserving both spatial details and spectral fidelity. Recent diffusion based methods exploit pretrained image priors by generating a low dimensional representation and subsequently mapping it to the hyperspectral domain. However, the observed panchromatic and hyperspectral images are typically imposed only through external reconstruction objectives, limiting their direct interaction with the diffusion prior. To address this issue, we propose dual-modality image-prompted diffusion model (DIDM) for zero shot hyperspectral pansharpening. DIDM encodes the low resolution hyperspectral and panchromatic observations into spectral and spatial prompt tokens, respectively, and injects them into intermediate features of a frozen remote sensing diffusion model through cross attention, allowing complementary spectral and spatial information to directly guide diffusion feature evolution. In addition, we introduce a panchromatic guided weighted pixel aware total variation regularizer that combines low resolution hyperspectral degradation fidelity and panchromatic response fidelity with gradient adaptive structural regularization, thereby preserving structural discontinuities while suppressing spurious variations in homogeneous regions. Extensive experiments on Pavia, Chikusei, and Houston under reduced resolution protocols show that DIDM achieves the best performance across all evaluated metrics, while full resolution evaluation on FR1 yields the highest HQNR among the compared methods. These results demonstrate that internal dual modality prompting and panchromatic guided structural regularization provide an effective balance between spatial detail enhancement and spectral preservation.