CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

📅 2026-04-29
📈 Citations: 0
Influential: 0
📄 PDF

career value

188K/year
🤖 AI Summary
Current vision–language models struggle to accurately capture radiologists’ clinical reasoning and visual attention mechanisms, resulting in chest X-ray interpretations that lack both accuracy and interpretability. This work introduces a large-scale multimodal dataset comprising over 100,000 clinical chains-of-thought and more than 6.6 million synchronized visual attention annotations from 501 radiologists across 71 countries, covering over 50,000 chest radiographs. For the first time, this resource systematically records and structures expert cognitive processes and visual search strategies. Models trained on this dataset significantly outperform existing approaches in pathology classification, visual faithfulness, temporal reasoning, and hallucination suppression. Moreover, they can predict diagnostic discrepancies between human–human and human–AI pairs, thereby enhancing diagnostic transparency and reliability.
📝 Abstract
Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current vision--language models are primarily trained on datasets of paired images and reports, not the cognitive processes and visual attention that underlie clinical reasoning. Here, we present CheXthought, a global, multimodal resource containing 103,592 chain-of-thought reasoning traces and 6,609,082 synchronized visual attention annotations across 50,312 multi-read chest X-rays from 501 radiologists in 71 countries. Our analysis reveals clinical reasoning patterns in how experts deploy distinct visual search strategies, integrate clinical context, and communicate uncertainty. We demonstrate the clinical utility of CheXthought across four dimensions. First, CheXthought reasoning significantly outperforms state--of--the--art vision--language model chain-of-thought in factual accuracy and spatial grounding. Second, visual attention data used as an inference--time hint recovers missed findings and significantly reduces hallucinations. Third, models trained on CheXthought data achieve significantly stronger pathology classification, visual faithfulness, temporal reasoning and uncertainty communication. Fourth, leveraging CheXthought's multi-reader annotations, we predict both human--human and human--AI disagreement directly from an image, enabling transparent communication of case difficulty, uncertainty and model reliability. These findings establish CheXthought as a resource for advancing multimodal clinical reasoning and the development of more transparent, interpretable vision--language models.
Problem

Research questions and friction points this paper is trying to address.

chest X-ray interpretation
clinical reasoning
visual attention
vision-language models
multimodal dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

chain-of-thought reasoning
visual attention
multimodal clinical reasoning
vision-language models
uncertainty communication
🔎 Similar Papers
No similar papers found.
S
Sonali Sharma
Center for Artificial Intelligence in Medicine and Imaging, Stanford University, Stanford, CA, USA
J
Jin Long
Department of Pediatrics, Stanford University School of Medicine, Stanford, California, USA
G
George Shih
Department of Radiology, Weill Cornell Medicine, New York, NY, USA
S
Sarah Eid
Department of Radiology, Brigham and Women’s Hospital, Boston, MA, USA
Christian Bluethgen
Christian Bluethgen
Radiologist, Clinician Scientist, USZ Zurich, AIMI Center, Stanford University
RadiologyThoracic ImagingMultimodal Machine Learning
F
Francine L. Jacobson
Department of Radiology, Brigham and Women’s Hospital, Boston, MA, USA
E
Emily B. Tsai
Department of Radiology, Stanford University, Stanford, CA, USA
G
Global Radiology Consortium
Global Radiology Consortium
Ahmed M. Alaa
Ahmed M. Alaa
Assistant Professor, UC Berkeley and UCSF
Machine LearningArtificial IntelligenceCausal InferenceAI for MedicineHealthcare
Curtis P. Langlotz
Curtis P. Langlotz
Professor of Radiology, Medicine, and Biomedical Data Science, Stanford University
machine learningcomputer visionnatural language processingdecision support systemstechnology assessment