InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过InfraBench评估AI代理在复杂基础设施管理任务中的表现,发现即使最强的代理也无法完美完成所有任务。
📝 Abstract
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.
Problem

Research questions and friction points this paper is trying to address.

computing infrastructure
AI agents
automation
complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

InfraBench
AI agents
infrastructure management
risk assessment
benchmark suite
🔎 Similar Papers
No similar papers found.