Measuring the Creativity of Frontier LLMs in Automated Research

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一套评估前沿LLM在自动化研究中创意性的指标,从价值性和新颖性两个维度进行评价,并发现变量级新颖性与研究表现最相关。
📝 Abstract
Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In this paper, we propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether a previously unexplored variable or variable combination is explored (Variable-level P-Novelty), and whether the idea directly follows retrieved external knowledge or departs from it (H-Novelty). Our evaluation shows that the models achieve relatively similar scores on most creativity metrics, but differ substantially in Variable-level P-Novelty, which reflects the breadth of research-space exploration. Further correlation and idea-level performance analyses show that Variable-level P-Novelty is the creativity dimension most strongly associated with research performance.
Problem

Research questions and friction points this paper is trying to address.

creativity
frontier LLMs
automated research
valueness
novelty
Innovation

Methods, ideas, or system contributions that make the work stand out.

creativity metrics
valueness
novelty
Variable-level P-Novelty
research performance
Y
Yiheng Zhao
Concordia University, Montreal, Canada
M
Mengzhuo Chen
Independent Researcher
Chengming Hu
Chengming Hu
McGill University
AI for Energy/TelecomAI-IoTCyber-Physical Security
P
Pengyi Liao
Concordia University, Montreal, Canada
Y
Yiran Pang
Florida Atlantic University, Boca Raton, USA