Selection-Aware Stress Testing for Interactive Agents

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决交互式代理评估中选择偏差问题,提出Selection-Aware Semantic Stress Testing方法,通过任务重加权和独立验证任务来提高评估的可靠性和稳定性。
📝 Abstract
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $τ$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.
Problem

Research questions and friction points this paper is trying to address.

Selection Bias
Interactive Agents
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selection-Aware
Semantic Stress Testing
Task Reweighting
Pre-execution Features
Conditional Asymptotic Validity
🔎 Similar Papers
No similar papers found.