Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了过程追踪对大型语言模型监督者决策标准的影响,使用信号检测理论分析五种模型在19个任务中的表现,发现详细的过程追踪会增加错误警报。
📝 Abstract
Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 19 compliance tasks (4,551 analyzed judgments), varying only trace detail and evidence labeling. With disconfirming evidence always visible, error detection remains near ceiling. Instead, elaborate traces shift the decision criterion toward rejection, increasing false alarms in susceptible overseers. Without option labels, human-validated reason coding shows about 60% of false alarms cite an inability to tie evidence to its option. Labels eliminate this stated reason, yet residual rejection of correct work persists in those overseers and rises with trace detail. Procedural traces thus act as governance artifacts that shape oversight decisions. AI auditors should be evaluated by their decision criterion and false-alarm behavior, alongside accuracy.
Problem

Research questions and friction points this paper is trying to address.

procedural traces
decision criterion
false alarms
oversight
large language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

procedural traces
decision criterion
false alarms
signal detection theory
oversight decisions
🔎 Similar Papers
No similar papers found.
Z
Zihan Chen
Stevens Institute of Technology
D
Di Zhu
Stevens Institute of Technology
L
Lei Zheng
University of Massachusetts Boston
W
Weiling Li
Stony Brook University