Locating Hidden Failures Makes Long-Horizon Agents More Reliable

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析2518个长周期AI代理轨迹,分类6967个错误为78种类型,使用名为Scout的40亿参数验证器来定位失败点,以提高代理任务成功率和可靠性。
📝 Abstract
As AI agents take on long, autonomous tasks, we increasingly oversee rather than perform the work, yet we still judge them almost entirely by whether they finally succeed. An outcome cannot reveal where a run went wrong, whether the agent recovered, or the irreversible harm it caused along the way, and where long-horizon agents fail remains unmapped. We study $2518$ agent trajectories across software engineering, computer use, and science, close to real deployment, and classify $6967$ mistakes into $78$ failure types. Failure follows a recurring signature: after its first mistake an agent often fails to recover and rarely catches the error itself, so the run continues unchecked while still looking correct; whether an agent recovers depends on the task and the environment's feedback, not on the agent framework running it. Long-horizon agents can do real harm on the way to a passing result: even runs scored as solved delete data, corrupt systems, or fabricate success rather than earning it. We release these human-verified annotations as Traverse, a benchmark on which six frontier judges struggle to locate failure regardless of scale: even the strongest correctly identifies the first mistake in fewer than a third of runs. Yet Scout, a $4$B verifier we trained, locates failure far better than these judges and transfers to domains it never saw. Used at test time to select among an agent's candidate runs, it raises task success above the agent's own single-attempt performance, without retraining the agent. By making failure cheap to locate and correct, this work is a foundation for more trustworthy long-horizon agents that learn from their own mistakes, and a practical path to overseeing increasingly autonomous AI.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agents
hidden failures
reliable AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-horizon Agents
Failure Localization
Scout Verifier
Traverse Benchmark
🔎 Similar Papers
No similar papers found.
Salman Rahman
Salman Rahman
University of California Los Angeles
Machine LearningNatural Language ProcessingLanguage Modeling
Yubin Kim
Yubin Kim
MIT
Health AIAI SafetyAgents
Mihir Parmar
Mihir Parmar
Research Scientist, Google
Large Language ModelsInstruction-TuningReasoning/PlanningMulti-Agent SystemsBio/Clinical NLP
A
A. Ali Heydari
Google Research
Genglin Liu
Genglin Liu
University of California, Los Angeles
Natural Language Processing
Simon A. Lee
Simon A. Lee
Ph.D. Student, UCLA
AI in HealthcareMachine LearningFoundation Models
Weizhi Zhang
Weizhi Zhang
University of Illinois Chicago
PersonalizationLarge Language ModelsAgents
Arian Hosseini
Arian Hosseini
Research Scientist, Google DeepMind, Mila
ReasoningAlignmentPlanningGeneralization
A
Ahmed A. Metwally
Google Research
Y
Yuzhe Yang
Google Research
Baharan Mirzasoleiman
Baharan Mirzasoleiman
UCLA
Machine LearningOptimizationSubmodularityML SustainabilityData-quality
Xin Liu
Xin Liu
Google
Computer Networks and Distributed Systems
Pavel Izmailov
Pavel Izmailov
Anthropic; NYU
Machine LearningDeep LearningLanguage ModelsReasoningAI Alignment
S
Saadia Gabriel
University of California, Los Angeles
M
Mark Malhotra
Google Research
Shwetak Patel
Shwetak Patel
University of Washington, Washington Research Foundation Endowed Professor, Computer Science
Ubiquitous ComputingHuman-Computer InteractionSensorsEmbedded Systems
Daniel McDuff
Daniel McDuff
Google and University of Washington
Affective ComputingDeep LearningHuman-Computer InteractionHuman-Centered AIComputer Vision
Hamid Palangi
Hamid Palangi
Google and University of Washington
Artificial IntelligenceMachine LearningNatural Language Processing