Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建Sci-MMR基准,解决多模态研究代理在多步骤证据支持的科学推理中的准确性和证据获取、整合问题。
📝 Abstract
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents
Problem

Research questions and friction points this paper is trying to address.

multi-step evidence-grounded reasoning
multimodal benchmarks
scientific evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sci-MMR
multi-step evidence-grounded reasoning
evidence acquisition
evidence integration
multimodal agents
💼 Related Jobs
No related jobs found.
J
Jiaqiang Li
Fudan NLP Group
Y
Yajie Yang
Fudan NLP Group
Zhiheng Xi
Zhiheng Xi
Fudan University
LLM ReasoningLLM-based Agents
J
Jiadong Chen
Fudan NLP Group
E
Enyu Zhou
Fudan NLP Group
Senjie Jin
Senjie Jin
Fudan University
natural language processing
Y
Yang Nan
Fudan NLP Group
Jiazheng Zhang
Jiazheng Zhang
Fudan University
Large Language ModelNatural Language ProcessingData Mining
H
Han Wang
Fudan NLP Group
Y
Yanxin Li
Fudan NLP Group
D
Dingwei Zhu
Fudan N LP Group
B
Bicheng Deng
Fudan NLP Group
Y
Yuhui Wang
Fudan NLP Group
Xiang Zheng
Xiang Zheng
Department of Computer Science, City University of Hong Kong
Reinforcement LearningTrustworthy AIEmbodied AI
Q
Qi Zhang
Fudan NLP Group
Lei Bai
Lei Bai
Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
Xingjun Ma
Xingjun Ma
Fudan University
Trustworthy AIMultimodal AIGenerative AIEmbodied AI
T
Tao Gui
Fudan NLP Group