MIRA: Medical Image Reflection for Agentic Diagnosis

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing medical vision-language agents, which often invoke diagnostic tools indiscriminately, thereby introducing noisy or misleading evidence due to a lack of validation regarding tool necessity and evidential consistency. To overcome this, the authors propose a novel diagnostic framework endowed with autonomous evidence-seeking and reflective verification capabilities. The framework dynamically orchestrates image processing and web-search tools while evaluating the relevance and consistency of retrieved evidence. A two-stage training strategy is employed: first, high-quality fine-tuning trajectories are generated via tool-augmented Monte Carlo tree search; second, an online reflection-based evolutionary mechanism refines decision-making through self-correction and adaptation. Evaluated across nine medical visual reasoning benchmarks, the method achieves an average score of 64.73—outperforming Qwen3-VL-8B by 7.44 points—with a valid tool invocation rate of 73.8% and a harmful judgment rate reduced to 1.6%.
📝 Abstract
Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/
Problem

Research questions and friction points this paper is trying to address.

medical visual agents
tool use
evidence verification
diagnostic reasoning
noisy evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

reflective verification
tool-augmented reasoning
medical visual agents
Monte Carlo Tree Search
reinforcement learning
💼 Related Jobs
No related jobs found.
S
Shengzhi Wang
Tongji University
J
Jun Yang
École Normale Supérieure – PSL
K
Kai Wu
Tongji University
Xiaozhong Ji
Xiaozhong Ji
Nanjing University
computer visionimage processingsuper resolution
Y
Yiwen Ye
Northwestern Polytechnical University
Ziyang Chen
Ziyang Chen
Peking University
Quantum key distributionQuantum random number generation
M
Mingliang Xiong
Tongji University
W
Wen Fang
Tongji University
M
Mingqing Liu
Tongji University
M
Mengyuan Xu
Tongji University
M
Miaoxuan Shan
People’s Public Security University of China
C
Caiyan Liu
Tongji University
B
Bin He
Tongji University
Qingwen Liu
Qingwen Liu
Tongji University
Wireless NetworkingAI