StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
StarHarness通过分层搜索进化环境特定的代理框架,解决模型与企业环境不匹配问题,提高基准性能20-35个百分点。
📝 Abstract
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fixed. The evolved harness can include prompt and task framing, tool interfaces, skills, MCP-backed providers, subagent structure, and agent-loop configuration. StarHarness constructs a compact evolution pool by stratifying tasks according to baseline failure behavior, separates proposer-visible search tasks from proposer-hidden selection tasks, and reserves held-out tasks for evaluating generalization. Across ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance, harness evolution improves full-benchmark performance by 20-35 percentage points over the default harness after 4-12 accepted changes per environment. These gains persist on tasks excluded from evolution and transfer without re-evolution across GPT and Qwen model families. Trace analysis links the improvements to interface repairs, environment conventions, and operational knowledge that compresses search, with fewer false-positive diagnoses and shorter trajectories in several settings. StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.
Problem

Research questions and friction points this paper is trying to address.

environment-specific agent harnesses
fixed model weights
enterprise environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stratified Search
Harness Evolution
Model-Environment Mismatch
Enterprise Environments
🔎 Similar Papers
No similar papers found.