Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing text-guided medical image segmentation methods that overlook frequency-domain information, leading to imprecise characterization of lesion textures and boundaries. To overcome this, we propose the first dual-domain cross-modal decoding framework that jointly leverages spatial and frequency domains. In the spatial domain, a Text-Guided Spatial Cross-Attention (TGSA) module aligns multi-scale visual features with clinical text, while in the frequency domain, a Spectral-Text Adaptive Modulation (STAM) mechanism enables frequency-aware channel recalibration. A lightweight two-stage refinement module further recovers high-resolution segmentation masks. By integrating clinical text guidance into both domains simultaneously, our approach achieves complementary linguistic supervision. Experimental results show state-of-the-art performance, with 91.46% Dice / 84.26% mIoU on QaTa-COV19 and 81.95% Dice / 69.42% mIoU on MosMedData+, surpassing the strongest baseline by +1.96 in Dice and +2.67 in mIoU on average.
📝 Abstract
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.
Problem

Research questions and friction points this paper is trying to address.

medical image segmentation
clinical text guidance
frequency domain
spatial alignment
cross-modal decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Domain Cross-Modal Decoding
Text-Guided Spatial Cross-Attention
Spectral-Text Adaptive Modulation
Frequency-Aware Decoding
Clinical Text-Guided Segmentation
💼 Related Jobs
No related jobs found.