Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Vision Transformers (ViTs) have shown immense potential in medical image analysis. However, standard pre-training via global image classification suffers from spatial collapse, where models rely heavily on background shortcuts rather than localising critical foreground lesions. To overcome this limitation and align visual evidence with precise medical semantics, we systematically investigate alternative pre-training paradigms.Specifically, we evaluate three independent forms of structured supervision: topological priors via graph self-supervision, dense pixel-level constraints via segmentation, and cross-modal semantic grounding via image-text pairs. Notably, our empirical analysis reveals that while all three forms of structured supervision successfully alleviate the global pooling bottleneck and steer visual attention towards foreground regions, image-text alignment achieves the most superior performance. By embedding high-dimensional diagnostic logic, the cross-modal approach not only anchors attention on precise visual evidence but also enables profound abstract reasoning. Extensive experiments demonstrate that this semantically enriched pre-training fundamentally enhances the model's feature representation. Consequently, when fine-tuned for downstream clinical classification tasks, our models achieve superior accuracy and yield highly interpretable attention maps focused on true pathological features, vastly outperforming vanilla classification baselines.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
spatial collapse
structured supervision
medical semantics
image-text alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured supervision
image-text alignment
semantic grounding
cross-modal approach
visual transformers
💼 Related Jobs
No related jobs found.
H
Hexiang Bai
Zhejiang University
Hanyang Xu
Hanyang Xu
Zhejiang University
X
Xiaoxue Li
Henan Normal University
X
Xiaoliang Wu
University of Southampton
Shangde Gao
Shangde Gao
Zhejiang University
Knowledge DistillationKnowledge AmalgamationAI for Science
Hongxia Xu
Hongxia Xu
Zhejiang University
AI4ScienceNanomedicineMedical imaging
K
Ke Liu
Zhejiang University; School of Software Technology, Zhejiang University