Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Existing approaches struggle to scalably evaluate the quality of large language models’ judgments on scientific ideas. This work proposes the Prediction-over-Time (PoT) benchmark framework, which constructs offline sandboxes via temporal slicing and predicts subsequent observable signals—such as citation counts or shifts in research agendas—based solely on evidence available up to a frozen cutoff date, thereby enabling verifiable evaluation without extensive expert annotations. Integrating tool-augmented agents, prompt ablation, and budget-scaling strategies, PoT is validated across over 30,000 instances spanning four scientific domains. Results show that increasing interaction budgets generally improves performance over non-agent baselines, while the efficacy of tool use is highly task-dependent. The framework further supports analyses of human–agent judgment alignment and enables controlled evaluations of agent-based scientific reviewing.