LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of failure attribution in long-horizon agents by constructing the first injection-free benchmark comprising 1,140 real-world trajectories. To facilitate diagnosis without additional training, we propose RCTA, a training-free root cause tracing method that leverages segment summary retrieval and instruction backtracking mechanisms. Experimental results demonstrate that RCTA achieves a responsible role localization accuracy of 51.1% and a root cause step precision of 24.1%. By filling the gap in benchmarks for real-world long-horizon agent failure attribution, this work provides an effective, training-free solution to enhance the interpretability and reliability of complex agent systems.
📝 Abstract
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon agent
Failure diagnosis
Root cause analysis
Responsible role attribution
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

LongRCA Bench
Root-Cause Trajectory Attribution
Training-free
Long-Horizon Agent
Failure Diagnosis
🔎 Similar Papers
No similar papers found.
Y
Yunfei Zhang
Computer Network Information Center, Chinese Academy of Sciences
B
Boyu Feng
Chongqing University
C
Changhua Pei
Computer Network Information Center, Chinese Academy of Sciences
Z
Zexin Wang
Computer Network Information Center, Chinese Academy of Sciences
Z
Zhihuang Peng
Computer Network Information Center, Chinese Academy of Sciences
X
Xinlong Liu
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences
H
Hengyue Jiang
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences
D
Difeng Ma
Computer Network Information Center, Chinese Academy of Sciences
J
Jiayi Zhang
Computer Network Information Center, Chinese Academy of Sciences
Y
Yongzhou Yao
Institute of Computing Technology, Chinese Academy of Sciences
Yanan Zhao
Yanan Zhao
NTU - NANYANG TECHNOLOGICAL UNIVERSITY
signal and information processinggraph generationdiffusion model
Fei Sun
Fei Sun
Associate Professor, Institute of Computing Technology, CAS; Alibaba
Natural Language ProcessingRecommender SystemAI Safety
Yintong Huo
Yintong Huo
Singapore Management University
AI4SEAIOpsLog analysisMLLM for SE
Zhaoyang Liu
Zhaoyang Liu
Tongyi Lab, Alibaba Group
LLMRecommendation
J
Jingjing Li
Computer Network Information Center, Chinese Academy of Sciences
G
Gaogang Xie
Computer Network Information Center, Chinese Academy of Sciences
Dan Pei
Dan Pei
Associate Professor of Computer Science, Tsinghua University
AIOpsTime Series Intelligence