What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为评估语言模型提出的研究想法的质量,本文引入了Lit2Test基准,通过可证伪的六领域契约来判断研究提案的质量,并对四个前沿模型进行了1200次盲评比较。
📝 Abstract
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.
Problem

Research questions and friction points this paper is trying to address.

Large language models
Research ideas
Judging criteria
Falsifiable outcomes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lit2Test
falsifiable outcome
benchmarking
reliability audit
research ideation
💼 Related Jobs
No related jobs found.
Z
Ziyue Wang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
A
Aomufei Yuan
Peking University
Y
Yiran Yao
Tianjin University
Linli Yao
Linli Yao
Peking University
multi-modal semantic understanding
H
Hongyao Zuo
Tianjin University
Z
Ziwen Gong
Hainan University
Yuanxin Liu
Yuanxin Liu
Peking University
Natural Language Processing
S
Shicheng Li
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Y
Yishuo Cai
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookupHash Table
Xu Sun
Xu Sun
Peking University
natural language processingdeep learningnatural language generationmulti-modal NLP
X
Xiaohui Li
Huawei Technologies
Haoli Bai
Haoli Bai
Huawei Technologies
natural language processingmodel compression