Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses whether authentic research ideas can be recovered solely from pre-publication literature by constructing a rigorously leakage-proof blind evaluation benchmark. We propose a reference-driven multi-agent tournament screening framework that integrates anonymous citation, cross-review, and collaborative mechanisms to effectively prevent data leakage while accurately reconstructing research trajectories. Experimental results demonstrate that this approach increases the idea matching rate by 23%–42%, achieving approximately 2.4 times the performance of the best single-model baseline. These findings validate the significant advantages and innovative value of multi-agent collaboration in academic idea tracing tasks, highlighting its potential for robust scholarly reconstruction without access to post-publication metadata or privileged information.
📝 Abstract
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.
Problem

Research questions and friction points this paper is trying to address.

Research Idea Recovery
Blind Benchmark
Pre-Publication Bibliography
Anti-Leakage Protocol
Innovation

Methods, ideas, or system contributions that make the work stand out.

Blind Benchmark
Anti-leakage Protocol
Multi-agent Pipeline
Cross-model Review
Swiss Tournament
💼 Related Jobs
No related jobs found.
S
Shaolong Chen
Prentis AI, San Francisco, CA
Y
Yanlin Fei
Prentis AI, San Francisco, CA
N
Nazhou Liu
Prentis AI, San Francisco, CA
X
Xinmiao Yu
Prentis AI, San Francisco, CA
L
Lei Li
Prentis AI, San Francisco, CA
Rahul Thapa
Rahul Thapa
Graduate Student, Stanford University
Machine LearningHealthcare AIData Science
M
Madalina Ciobanu
Prentis AI, San Francisco, CA
Q
Qingqing Mao
Prentis AI, San Francisco, CA; Titan Holdings, San Francisco, CA
R
Ritankar Das
Prentis AI, San Francisco, CA; Titan Holdings, San Francisco, CA