Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CAST框架,通过自动发现解剖区域和对比解码过程,在无需标注的情况下减少医学视觉语言模型中的幻觉问题。
📝 Abstract
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.
Problem

Research questions and friction points this paper is trying to address.

Medical Vision-Language Models
Hallucination
Anatomical Awareness
Ground Truth Annotations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Intervention
Anatomical Awareness
Contrastive Decoding
Automatic Region Selection
💼 Related Jobs
No related jobs found.
Y
Yifan Lu
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE; MedOS, Abu Dhabi, UAE
A
Adinath Dukre
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE; MedOS, Abu Dhabi, UAE
A
Abhijit Das
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE; MedOS, Abu Dhabi, UAE
Z
Ziyun Zou
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE; MedOS, Abu Dhabi, UAE
Haolin Yang
Haolin Yang
University of Chicago
large language modelsnatural language processing
Yutong Xie
Yutong Xie
Assistant Professor, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Medical image analysisComputer visionDeep learningMulti-modal learning
Imran Razzak
Imran Razzak
MBZUAI, Abu Dhabi
Human-Centered AIMedical Image AnalysisMedical Artificial IntelligenceComputational Biology