SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SciRIGOR框架,通过评估科学编码代理生成的代码、结果和声明之间的支持关系,解决仅依赖最终输出无法验证结论科学性的问题。
📝 Abstract
Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations. On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8% versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain success. These findings show that internal coherence does not establish scientific correctness: evaluation must verify support along the complete data-to-claim path.
Problem

Research questions and friction points this paper is trying to address.

scientific analysis
evidence-grounded
claim-support
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence-grounded multimodal scientific analysis
SciRIGOR
typed evidence graphs
claim-support paths
🔎 Similar Papers
No similar papers found.