AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有评估方法不能准确衡量AI科学家科学构思能力的问题,本文提出AgentIdeaBench基准,通过静态观察和主动探索两种设置来评价33个语言模型在科学构思上的表现。
📝 Abstract
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.
Problem

Research questions and friction points this paper is trying to address.

Scientific Ideation
Autonomous AI Scientists
Static Observation
Active Exploration
Retrieval-and-Reasoning Workflow
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Ideation
Active Exploration
Static Observation
Scientific World Modeling
Capability Gating
💼 Related Jobs
No related jobs found.
Y
Yunxiang Mo
Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China
Tianshi Zheng
Tianshi Zheng
HKUST
Natural Language ProcessingLogical InferenceScientific DiscoveryResearch Agent
Yisen Gao
Yisen Gao
The Hong Kong University of Science and Technology
Geometric deep learningGenerative modelNeural graph database
R
Rui Wang
Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China
N
Newt Nguyen Kim Hue Nam
Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China
K
Kelvin Kiu Wai Tam
Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China
Jiaxin Bai
Jiaxin Bai
Hong Kong University of Science and Technology
Natual Language Processing
Yangqiu Song
Yangqiu Song
HKUST
Artificial IntelligenceData MiningNatural Language ProcessingKnowledge GraphsCommonsense Reasoning
G
Ginny Wong
NVIDIA AI Technology Center (NV AITC), NVIDIA, Santa Clara, USA
Simon See
Simon See
nvidia
applied mathematicsAImachine learningHigh Performance ComputingSimulation