BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出BenchShield,通过模型支持的检测方法解决LLM代理评估中的奖励篡改问题,提高检测准确率并降低成本。
📝 Abstract
LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not provide reusable evidence that a concrete run remained within its intended evaluation boundary. This paper presents BenchShield, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation. BenchShield grounds detection in a finite lifecycle model of an evaluation's reward-relevant events. Within the benchmark infrastructure, two complementary analyses operate over this model. A static, phase-aware taint analysis exposes reward-hacking paths before a run. Its runtime counterpart uses infrastructure-side evidence to attribute concrete agent use and emit evidence-backed claims. We construct BenchShield Trajectories, a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks. Compared with an agentic hackability scanner baseline on the same tasks and model, BenchShield improves full-chain recall from 23-94% to 77-100%, same-vector coverage from 16-56% to 43-78%, and reduces per-task cost by up to 65%. Its runtime analysis achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Problem

Research questions and friction points this paper is trying to address.

reward hacking
LLM-agent evaluation
interactive evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

BenchShield
reward integrity
formal model
taint analysis
infrastructure-side evidence
💼 Related Jobs
No related jobs found.
S
Shenghan Zheng
Dartmouth College
Z
Zonglin Di
Independent
Y
Yimin Liu
Ohio State University
K
Kyoung Whan Choe
RLWRLD
J
Jiankai Sun
Independent
H
Heguang Lin
The Scripps Research Institute
P
Penghao Jiang
University of New South Wales
Y
Yifeng He
University of California, Davis
X
Xiao Cheng
Macquarie University
J
Jicheng Wang
University of California, Davis
W
Wenbo Chen
Amazon
A
Alex Yates
Independent
Y
Yinzhe Zhao
Independent
B
Bingran You
BenchFlow
Yuan Gao
Yuan Gao
Department of Applied Mathematics, University of Washington
OptimizationMachine LearningControlBioinformatics
A
Ayush Munot
Independent
Shubham Gaur
Shubham Gaur
University of California Santa Cruz
Machine LearningNatural Language ProcessingAlignmentEfficiencyMultimodality
Z
Zhe Ye
UC Berkeley
H
Hao Wang
UC Berkeley
X
Xiangyi Li
BenchFlow
Dawn Song
Dawn Song
Professor of Computer Science, UC Berkeley
Computer Security and Privacy
Christophe Hauser
Christophe Hauser
Dartmouth College
Systems securitybinary analysisdirtbikes