DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning
This work addresses the limitations of existing remote sensing image change captioning methods, which rely on a single autoregressive generation paradigm and tend to favor frequent vocabulary while overlooking discriminative differences. To overcome this, the authors propose a difference-aware feature modeling framework that leverages a text-guided gated contrastive loss to steer the visual encoder—via linguistic cues—toward salient change regions. The approach integrates a pretrained change detection model with a multi-scale joint feature modeling module to comprehensively capture spatiotemporal discrepancies between multi-temporal images. By transcending the constraints of conventional generative paradigms, the method achieves significant improvements in both caption accuracy and discriminative expression across multiple remote sensing change captioning benchmarks, demonstrating its effectiveness and novelty.