Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自动科研代理在开放性研究任务中可能错过重要分析、使用不当方法等问题,本文提出AutoSciRub框架,通过生成特定任务的评价标准来指导执行和迭代修订。
📝 Abstract
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).
Problem

Research questions and friction points this paper is trying to address.

autonomous scientific research agents
open-ended research tasks
evaluation criteria
task specification
evidence-based conclusions
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation-first
executable rubric
autonomous scientific research
iterative revision
criterion-level verification
💼 Related Jobs
No related jobs found.
Xuehai Wang
Xuehai Wang
Department of Learning, Informatics, Management & Ethics, Karolinska Institutet
Machine learningImmunologyBiomedical AIOncologyMultimodal
H
Haowei Qin
University of Electronic Science and Technology of China
T
Tongxin Liu
Beijing University of Posts and Telecommunications
J
Junkai Li
Zhejiang University of Technology
B
Buqiang Xu
Zhejiang University
Jintian Zhang
Jintian Zhang
Zhejiang University
NLPLLMs
Y
Yijun Chen
Zhejiang University
Z
Zirui Xue
Zhejiang University
Shumin Deng
Shumin Deng
National University of Singapore
NLPLLM Planning & ReasoningLLM AgentKGIE