Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning
This work addresses the limited clinical reasoning capability and poor interpretability of general-purpose vision-language models (VLMs) in radiological diagnosis. We propose a weakly supervised paradigm that requires no annotated image-lesion pairs, leveraging only free-text radiology reports. Our method automatically parses unstructured reports into structured, stepwise Chain-of-Thought (CoT) reasoning paths, then integrates contrastive image-report alignment with multi-granularity clinical reward-guided reinforcement fine-tuning. To our knowledge, this is the first framework to distill stepwise diagnostic supervision signals—aligned with radiologists’ cognitive reasoning—from raw text reports alone. Zero-shot evaluation on MIMIC-CXR demonstrates substantial improvements: +0.24 in disease classification AUC, +0.23 in lesion localization mIoU, and +0.22 in report generation BLEU score, outperforming state-of-the-art methods. The approach establishes a novel, interpretable, and scalable paradigm for training medical VLMs.