XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨使用大型语言模型(LLMs)作为评估可解释AI(XAI)解释质量的机制,通过引入XAI-Arena框架来解决现有方法依赖主观判断导致的问题,提供了一种可扩展且可重复的方法。
📝 Abstract
Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.
Problem

Research questions and friction points this paper is trying to address.

XAI
Explanation Quality
Subjective Human Judgment
Reproducibility
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

XAI-Arena
LLM-as-a-judge
scalable evaluation
reproducibility
stakeholder-sensitive