Explaining AI Agents Through Execution Traces

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种后验XAI框架,通过执行轨迹生成结构化报告和自然语言解释,解决了AI代理在多步骤决策中缺乏过程透明度的问题。
📝 Abstract
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems, motivating a paradigm shift toward approaches specifically designed for AI Agents. To address this gap, we present a post-hoc XAI framework that transforms a lengthy agent's execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior. Because it relies solely on execution traces, the framework applies across different agent architectures, environments, and tasks. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces high-quality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming naive LLM-generated explanations.
Problem

Research questions and friction points this paper is trying to address.

AI Agents
Explainable AI (XAI)
Execution Traces
Transparency
Post-hoc Explanation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Execution Traces
Post-hoc XAI Framework
Natural-language Explanation
Process-level Transparency
AI Agents
V
Vittoria Vineis
Sapienza University of Rome, Rome, Italy
F
Fabiano Veglianti
Sapienza University of Rome, Rome, Italy
L
Lorenzo Antonelli
Sapienza University of Rome, Rome, Italy
C
Claudia Di Carlo
Sapienza University of Rome, Rome, Italy
M
Matteo Silvestri
Sapienza University of Rome, Rome, Italy
Gabriele Tolomei
Gabriele Tolomei
Associate Professor of Computer Science at Sapienza University of Rome
Machine LearningExplainable AIFederated LearningAdversarial LearningWeb Search & Advertising