The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文针对代理基准测试中的双重测量混淆问题,通过将关键决策转移给模型、使用基于真实值的评分及报告更全面的可靠性指标来解决。
📝 Abstract
Agent benchmarks are increasingly used to compare large language models (LLMs) and guide deployment decisions, yet benchmark scores are meaningful only if they measure model capability rather than properties of the evaluation pipeline. We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, while the scorer evaluates outputs using criteria that may not reflect task correctness. We unify these issues within a measurement-theoretic framework that characterizes when benchmark scores can be interpreted as evidence of model capability, and instantiate it with an audit-and-repair protocol that (i) transfers execution-critical decisions from the scaffold to the model, (ii) replaces shape-based evaluation with seeded ground-truth scoring, and (iii) reports reliability beyond the mean through worst-case and tail-risk metrics. Experiments on ComtradeBench show that the joint intervention transforms a nearly flat leaderboard into a reliability spectrum that distinguishes both average performance and robustness across seeds. Applying the audit to existing benchmarks further shows that scorer validity is benchmark-specific, whereas scaffold ownership is an uncontrolled axis wherever we probed it. Our results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.
Problem

Research questions and friction points this paper is trying to address.

Double Measurement Confound
Agent Benchmarks
Scaffold
Ground-Truth Scoring
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

de-scaffolding
ground-truth scoring
reliability beyond the mean
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yonghong Zhang
Universidad Autónoma de Madrid, Spain
S
Shadi Motaali
Universidad Autónoma de Madrid, Spain
V
Vu Phong Dinh
IMDEA Nanociencia, Spain
A
Avin Piroutiniya
Universidad Complutense de Madrid, Spain
J
Jorge E. López de Vergara
Universidad Autónoma de Madrid, Spain
L
Luis de Pedro
Universidad Autónoma de Madrid, Spain
Ricardo Correia
Ricardo Correia
Universidad Autónoma de Madrid, Spain
I
Isabel M. Parra
Universidad Autónoma de Madrid, Spain
Y
Yong Xie
Spanish National Research Council (CSIC), Madrid, Spain