CARE: Confidence-Aware Reasoning for Reliable Medical VQA

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the critical misalignment between confidence and diagnostic accuracy in medical multimodal large language models for visual question answering, which undermines clinical trustworthiness. To resolve this, the authors propose the CARE framework, which explicitly integrates confidence calibration into the reinforcement learning reward signal—a first in this domain. The approach begins with supervised fine-tuning using Medical-CoT–synthesized data, followed by optimization via a novel Group Relative Policy Optimization (GRPO) algorithm incorporating a confidence-aware reward mechanism. Evaluated on three medical VQA benchmarks, CARE simultaneously achieves state-of-the-art diagnostic accuracy, the lowest expected calibration error, and the lowest hallucination rate, substantially enhancing both the reliability and correctness of model predictions.
📝 Abstract
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel $\textbf{Confidence-Aware Reward (CAR)}$ mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, $\textbf{CARE}$ achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.
Problem

Research questions and friction points this paper is trying to address.

confidence miscalibration
medical visual question answering
multimodal large language models
clinical trust
diagnostic accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Confidence-Aware Reasoning
Medical VQA
Reinforcement Fine-Tuning
Calibration
Chain-of-Thought
💼 Related Jobs
No related jobs found.