Evo-Bench: Can Language Models Improve Agent Harness?

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluation methods struggle to independently assess an agent’s ability to autonomously optimize its operational harness, often conflated by the underlying model’s performance and lacking characterization of long-term evolution. To address this, this work proposes Evo-Bench—the first benchmark specifically designed to evaluate language models’ capacity for self-evolving their harnesses—spanning search, office, and general domains. By leveraging a harness-guided construction framework, auxiliary task evolution, and sensitivity-aware hierarchical partitioning, Evo-Bench enables systematic and decoupled assessment of harness evolution capabilities. Experiments show that leading models achieve up to a 16.6-point improvement on Evo-Bench, approaching human-designed optimal baselines, with strong performance in general and search tasks, though gaps remain in office tasks requiring specific procedural knowledge. These results validate both the high transferability of synthesized harnesses and the effectiveness of the proposed evaluation framework.
📝 Abstract
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
Problem

Research questions and friction points this paper is trying to address.

harness evolution
autonomous agents
benchmarking
large language models
framework optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

harness evolution
Evo-Bench
autonomous agent optimization
sensitivity-aware stratified splitting
auxiliary-task evolution
💼 Related Jobs
No related jobs found.