ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

πŸ“… 2026-08-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the absence of effective benchmarks for evaluating the alignment, logical coherence, and evolutionary completeness of automated scientific research systems along human-like research trajectories. It introduces ARAC-Bench, the first evaluation framework that translates implicit reviewer expertise into staged, quantifiable scoring criteria. Centered on academic cognitive skills and a three-stage diagnostic protocol, ARAC-Bench emphasizes high-quality research processes rather than final outputs alone, establishing a modular and traceable tri-dimensional diagnostic system spanning proposal, experimentation, and synthesis. Evaluation across 11 state-of-the-art systems reveals a maximum alignment score of only 67.9/100, highlighting substantial gaps. The framework’s scores exhibit a strong correlation (r = 0.8141) with rankings of PhD candidates, validating its effectiveness in assessing research capability.
πŸ“ Abstract
The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.
Problem

Research questions and friction points this paper is trying to address.

Auto-Research
alignment
completeness
research trajectory
evaluation benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Auto-Research
ARAC-Bench
Researcher-Mimicking Evaluation
Academic Cognition Skills
Three-stage Diagnostic Protocol
J
Jiale Cui
School of Software Technology, Zhejiang University
Y
Yueyao Yuan
School of Software Technology, Zhejiang University
K
Kaixi Zhong
School of Software Technology, Zhejiang University
Xiaogang Xu
Xiaogang Xu
CUHK
Large ModelMulti-Modality AIAIGCGenerative PhotographyAI Security
J
Jiafei Wu
School of Software Technology, Zhejiang University
Zhe Liu
Zhe Liu
Professor, Zhejiang University
Cryptographic EngineeringComputer ArithmeticPost-Quantum Cryptography