Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大模型生成研究想法难以评估的问题,本文提出Ideation Arena平台,通过人类专家的两两比较评估方法来评价这些想法。
📝 Abstract
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideation Arena, a battle style platform that evaluates research ideas through pairwise human assessment. Ideation Arena evaluates ideas generated by 14 frontier LLMs and 5 research agent architectures built on 2 base models. To ensure a common starting point, Ideation Arena builds shared literature contexts from papers familiar to the participating researchers and provides the same contexts to all LLMs and agents. We collect over 6,000 double blind pairwise comparisons from 105 active computer science researchers and construct an Elo rating leaderboard of proposal-stage expert preferences in computer science under a shared closed-context protocol. We validate the rankings through interrater agreement and robustness analyses, showing that the leaderboard remains stable under changes in annotator composition and domain coverage. Our results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models. We further construct Ideation Arena Eval, a benchmark for assessing whether automated evaluators align with human preferences in research ideation. Experiments with current LLM judges show that they still cannot reliably reproduce expert preferences, with the best judge reaching 72.56% Soft Accuracy on Overall Quality. Our code, data, and leaderboards are available at https://github.com/foss12138/Research-Ideation-Arena.
Problem

Research questions and friction points this paper is trying to address.

Evaluating research ideas
Large Language Models (LLMs)
Scientific value
Objective criteria
Human assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elo rating leaderboard
pairwise human assessment
automated evaluators alignment
Zhiyu Chen
Zhiyu Chen
Amazon
Conversational AILarge Language ModelsInformation RetrievalNatural language Processing
K
Keyu Zhao
Tsinghua University
J
Jigao Fu
Zhongguancun Institute of Artificial Intelligence
D
Dong Liang
Zhongguancun Institute of Artificial Intelligence
Y
Yanbiao Wu
Zhongguancun Institute of Artificial Intelligence
Jiaoyang Li
Jiaoyang Li
Assistant Professor at Robotics Institute, Carnegie Mellon University
Artificial IntelligenceMulti-Agent/Robot SystemsHeuristic SearchAutomated Planning
H
Haidong Xue
Zhongguancun Institute of Artificial Intelligence
X
Xinhua Zeng
Fudan University
Y
Yuanyi Zhen
Zhongguancun Academy
Fengli Xu
Fengli Xu
Tsinghua University
LLM AgentData ScienceSocial ComputingScience of ScienceUrban Science
Yong Li
Yong Li
Professor, Electronic Engineering, Tsinghua University
Urban ScienceData MiningAI for Science