Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DualViewEval方法,通过联合利用结果和过程关系来压缩代理基准测试,有效减少了评估成本并提高了准确性。
📝 Abstract
Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall's $τ$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.
Problem

Research questions and friction points this paper is trying to address.

agent benchmarking
benchmark compression
final-score distributions
performance redundancy
trajectory analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

DualViewEval
agent benchmark compression
outcome and process relations
miniset learning
X
Xinshuai Guo
Hunyuan Team, Tencent; Tsinghua University
Junjie Wu
Junjie Wu
Center for High Pressure Science & Technology Advanced Research
Physics
D
Dolly Deng
Hunyuan Team, Tencent
Y
Yinghui Li
Tsinghua University
H
Hai-Tao Zheng
Tsinghua University
S
Suncong Zheng
Hunyuan Team, Tencent
M
Maxm Pan
Hunyuan Team, Tencent