The Answer Is Not the Argument

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在AI监管中提供参考答案是否有助于验证推理过程。通过对比有无答案情况下的监控效果,发现提供答案主要提高了结论一致性检查而非独立验证论证过程的有效性。
📝 Abstract
Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning verification or mainly exposes incorrect conclusions. We collected 237 step-numbered solutions to 79 Humanity's Last Exam physics questions from three frontier models, with no inserted errors, and independently labelled final-answer correctness and the first false step. The reference standard combined physicist annotations, an independent LLM debate, and source-masked adjudication. This yielded 24 critical traces in which the answer was correct but the trace contained a genuine error. 8 LLM monitors evaluated traces blind, with an unverified or certified answer, or after a blind commitment. Certification raised mean balanced accuracy from 0.637 to 0.796, while exact first-error localization rose from 0.261 to 0.379. Certification changed recall (the fraction of error traces flagged as erroneous) from 0.653 to 0.951 on wrong-answer traces but from 0.521 to 0.438 on critical traces; the contrast had the same direction for all 8 monitors (question-bootstrap 95% CI [+0.256, +0.506]). After blind commitment, monitors shown the answer newly flagged 93.8% of previously passed wrong-answer traces as erroneous, but only 18.0% of critical traces. Answer access therefore improves conclusion-consistency checking rather than independent verification of the supporting argument. For AI safety, these traces provide a benign analogue of reward hacking: an acceptable output does not establish that the process producing it was sound. Although the errors studied here were ordinary and mostly non-load-bearing rather than adversarial, trusted-answer evaluations may similarly overstate monitoring capability when acceptable outputs conceal unsound reasoning.
Problem

Research questions and friction points this paper is trying to address.

chain-of-thought monitoring
AI oversight
reasoning verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

chain-of-thought monitoring
reasoning verification
answer access
conclusion-consistency checking
🔎 Similar Papers
No similar papers found.
W
Will Yeadon
Department of Physics, Durham University, Durham DH1 3LE, UK
S
Sergio Juárez
Escuela de Ingeniería de Telecomunicación, Department of Signal Theory and Communications, University of Vigo, Vigo E-36310, Spain
P
Paul Mackay
Department of Physics, Durham University, Durham DH1 3LE, UK
T
T. J. Dowling
Independent Researcher
E
Elise Agra
Department of Physics, Durham University, Durham DH1 3LE, UK
O
Oto-obong Inyang
Department of Physics, Durham University, Durham DH1 3LE, UK
A
Arin Mizouri
Department of Physics, Durham University, Durham DH1 3LE, UK
C
Craig P. Testrow
Department of Physics, Durham University, Durham DH1 3LE, UK