🤖 AI Summary
Current XAI evaluation lacks standardized, reliable, and manipulation-resistant metrics—primarily due to the absence of ground-truth explanations, resulting in low credibility, high regulatory compliance risks, and difficulties in cross-method comparison. To address this, we propose a novel evaluation paradigm that is human-centered, scenario-adaptive, and manipulation-resistant, integrating human-subject experiments, domain-knowledge modeling, regulatory alignment analysis, and principled benchmark design. We develop domain-specific evaluation benchmarks for high-stakes applications—including healthcare and finance—and advance a multi-stakeholder standardization framework. Our approach shifts XAI evaluation from subjective, fragmented practices toward a systematic, trustworthy, compliant, and comparable methodology. The framework enables scalable, empirically grounded assessment aligned with real-world deployment requirements and regulatory expectations. (149 words)
📝 Abstract
This position paper emphasizes the critical gap in the evaluation of Explainable AI (XAI) due to the lack of standardized and reliable metrics, which diminishes its practical value, trustworthiness, and ability to meet regulatory requirements. Current evaluation methods are often fragmented, subjective, and biased, making them prone to manipulation and complicating the assessment of complex models. A central issue is the absence of a ground truth for explanations, complicating comparisons across various XAI approaches. To address these challenges, we advocate for widespread research into developing robust, context-sensitive evaluation metrics. These metrics should be resistant to manipulation, relevant to each use case, and based on human judgment and real-world applicability. We also recommend creating domain-specific evaluation benchmarks that align with the user and regulatory needs of sectors such as healthcare and finance. By encouraging collaboration among academia, industry, and regulators, we can create standards that balance flexibility and consistency, ensuring XAI explanations are meaningful, trustworthy, and compliant with evolving regulations.