Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

๐Ÿ“… 2026-09-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ไธบๅฎก่ฎกAI็ ”็ฉถไปฃ็†๏ผŒๆๅ‡บๅ‘็Žฐ่ฎค่ฏๅ่ฎฎ(DCP)๏ผŒ้€š่ฟ‡ๆ‰ง่กŒๆขๅคๅ’Œๅ้ฆˆๆต‹่ฏ•้ชŒ่ฏ็ป“ๆžœ็š„ๆœ‰ๆ•ˆๆ€งใ€‚
๐Ÿ“ Abstract
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.
Problem

Research questions and friction points this paper is trying to address.

AI research agents
result validation
Discovery Certification Protocol
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discovery Certification Protocol
recovery and feedback tests
useful improvement
registered starting information
finite-sample bound
๐Ÿ’ผ Related Jobs
No related jobs found.
J
Jingjie Ning
School of Computer Science, Carnegie Mellon University
Shanshan Zhong
Shanshan Zhong
Carnegie Mellon University
Language ModelsMultimodal UnderstandingMultimodal Generation
Xiaochuan Li
Xiaochuan Li
Carnegie Mellon University
Machine LearningNatural Language Processing
J
Ji Zeng
School of Computer Science, Carnegie Mellon University