Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation
To address coarse-grained multimodal alignment and insufficient feature fusion in referring segmentation of remote sensing images, this paper proposes a fine-grained cross-modal collaborative segmentation framework. The method introduces two key components: (1) a Correlation Fusion Module (CFM) that enables pixel-wise semantic alignment between textual and visual features via cross-modal correlation modeling; and (2) a Multi-Scale Refinement Convolution (MSRC) integrated with an adaptive noise-augmented Transformer-based visual encoder, which jointly captures multi-directional, multi-scale object structures and orientation invariance. Evaluated on the RRSIS-D benchmark, the proposed approach achieves significant improvements over existing state-of-the-art methods, attaining a 3.2% gain in mean Intersection-over-Union (mIoU). The source code is publicly available.