MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
This study addresses the underutilization of textual knowledge and the limitations of late-stage language guidance in medical image segmentation by proposing an end-to-end vision-language framework. Innovatively transforming text from a posterior condition into continuous structured supervision, the method employs bidirectional fusion and class/region-level multi-granularity concept alignment to integrate clinical text throughout the encoding process. This enables joint evolution of visual-linguistic representations for precise segmentation guidance. The proposed approach achieves state-of-the-art performance across CT/MR multi-organ, cardiac, and tumor segmentation benchmarks while effectively supporting real-world free-text supervision. Consequently, this work significantly enhances both clinical applicability and model generalizability by establishing text as a persistent supervisory signal rather than a mere post-hoc constraint.