Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过Evidence Decoupling Decoder方法分析文本在多模态医学图像分割中的作用,揭示了不同数据集中文本影响的差异。
📝 Abstract
Pretrained vision-language models (VLMs) have shown promising performance in medical image segmentation by incorporating clinical text. However, it remains unclear how much textual information actually contributes to pixel-level predictions. In this work, we systematically investigate the role of text in multimodal medical image segmentation. We first analyze several commonly used fusion strategies and find that segmentation performance is largely insensitive to the choice of fusion module. To further understand modality interactions, we propose an Evidence Decoupling Decoder (EDD) based on evidential deep learning and deep supervision. EDD serves as an internal representation analysis tool that decomposes image evidence and text-modulated evidence throughout the decoding process while maintaining competitive segmentation performance. Experimental results show that the sensitivity to text perturbation varies substantially across datasets. On BUSI and BTMRI, removing text causes catastrophic performance drops, indicating strong model reliance on textual input. On ISIC and Kvasir-SEG, text exerts relatively marginal influence. We further find that text affects predictions mainly through global semantic modulation rather than independent spatial localization, and that the specific semantic components driving text sensitivity differ across datasets. These findings provide a deeper understanding of modality interaction in multimodal medical image segmentation and offer practical insights for future model design.
Problem

Research questions and friction points this paper is trying to address.

text sensitivity
multimodal medical image segmentation
pixel-level predictions
modality interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence Decoupling Decoder
Evidential Deep Learning
Multimodal Medical Image Segmentation
Text Sensitivity
🔎 Similar Papers
Ziquan Liu
Ziquan Liu
Assistant Professor, Queen Mary University of London
machine learning
Z
Zhewei Zhu
School of Information and Control Engineering, Southwest University of Science and Technology, Mianyang 621010, China
X
Xuyang Shi
School of Information and Control Engineering, Southwest University of Science and Technology, Mianyang 621010, China