A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation
In medical image segmentation guided by text, conventional data augmentation (e.g., rotation, flipping) disrupts cross-modal spatial alignment between images and text, degrading performance. To address this, we propose an early-fusion framework that projects textual embeddings into the visual space *before* augmentation and synthesizes semantically consistent, interpretable pseudo-images via a lightweight generator—thereby bridging the modality gap while preserving spatial coherence. Our method is architecture-agnostic, requiring no modifications to downstream segmentation models. Evaluated across three medical segmentation tasks and four state-of-the-art segmentation frameworks, it achieves new SOTA results. Visualizations confirm that the generated pseudo-images accurately localize target anatomical regions. The core innovations are: (1) the “pre-augmentation multimodal fusion” paradigm, and (2) a text-driven mechanism for generating interpretable, semantically grounded pseudo-images.