DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.
Problem

Research questions and friction points this paper is trying to address.

large language models
general-purpose agents
evaluation benchmark
deployment decisions
workload variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deployment-Aware
Reliable Evaluation
Multi-Step Execution
Tool Use
🔎 Similar Papers
No similar papers found.
Y
Yu Liu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences; MiLM Plus, Xiaomi Inc.
Zhilin Liu
Zhilin Liu
School of Public Policy & Management, Tsinghua University
urban planning and policyurban governancehousing policysustainability
Zhiwei Yang
Zhiwei Yang
Guangzhou Institute of Technology, Xidian University, Guangzhou, China
Deep LearningComputer VisionAnomaly Detection
S
Shaojie Zhang
MiLM Plus, Xiaomi Inc.
Z
Zheyuan Deng
Department of Computer Science, Brown University
T
Tingwei Huang
MiLM Plus, Xiaomi Inc.
Zhenbo Luo
Zhenbo Luo
XiaoMi
Vision Language ModelComputer Vision
Lei Jiang
Lei Jiang
Technical Institute of Physics and Chemistry, Chinese Academy of Sciences
bio-inspired interfacial materials with superwettability
Y
Yanbing Liu
School of Cyber Security, University of Chinese Academy of Sciences
P
Pei Fu
MiLM Plus, Xiaomi Inc.