Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对LLM在科学实验中产生的方法幻觉问题,提出ABE-Ralph框架,通过结构化实验约束和多步骤验证来确保实验设计的忠实性和结果的有效性。
📝 Abstract
LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show that agents often produce methodological hallucinations: silently reducing datasets or training budgets, replacing failed learning or generative components with lookup or oracle functions, or drawing conclusions from resource-limited settings where a method's claimed advantage disappears. To detect these failures, we introduce ABE-Ralph, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification. Across 30 long-horizon reproduction runs covering 12 machine learning domains, ABE-Ralph achieves a 93% robust execution rate and identifies five scientific failure modes. In 23 NatureBench discovery tasks, ABE-Ralph matches or exceeds state-of-the-art performance on 5 tasks. These results show that reliable evaluation of AI scientists must assess whether the experimental design faithfully tests the intended claim and whether the resulting evidence supports it, rather than treating code execution or plausible metrics as evidence of scientific success.
Problem

Research questions and friction points this paper is trying to address.

LLM
methodological hallucinations
experimental fidelity
scientific research
reliable evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

ABE-Ralph
methodological hallucinations
structured experimental constraints
multi-level verification
experimental fidelity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lezhi Yu
College of Computer Science and Technology, Zhejiang University, Hangzhou, China
Xiaogang Xu
Xiaogang Xu
CUHK
Large ModelMulti-Modality AIAIGCGenerative PhotographyAI Security
Y
Yuhua Zhou
College of Computer Science and Technology, Zhejiang University, Hangzhou, China
Shuibing He
Shuibing He
Professor of Zhejiang University
Intelligent ComputingStorage SystemsProcessing-in-MemoryComputer Architecture
A
Aimin Pan
Zhejiang Lab, Hangzhou, China