Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决科学论文实验再现难题,提出AgentActionBench框架,通过记录和评估基于代理的实验再现过程,涵盖机器学习和AI4Science领域。
📝 Abstract
Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers, including 120 ML papers and 30 AI4Science papers. A human-annotated subset covering 10% of the benchmark provides validation data, while model-assisted augmentation expands the full benchmark to more than 10,000 rubric items. Experimental results show that current systems remain limited, with execution as the primary bottleneck. Meanwhile, the strong Pearson and Spearman correlations between model-generated and human-annotated rubrics validate the reliability of our scalable rubric-generation approach.
Problem

Research questions and friction points this paper is trying to address.

Reproducibility
Scientific Progress
Large Language Model (LLM) Agents
Machine Learning (ML)
AI4Science
Innovation

Methods, ideas, or system contributions that make the work stand out.

AgentActionBench
MCP-based Action Recorder
process-oriented benchmark
AI4Science
scalable rubric-generation