🤖 AI Summary
This study investigates whether users can distinguish between valid generalization and erroneous representation in conceptual explainable AI (C-XAI) within railway safety assessment tasks. Adopting an experimental psychology paradigm, we developed a simulated C-XAI system that presents hazard judgments alongside concept-based explanations via image patches, systematically manipulating concept relevance (e.g., track relations vs. actions) and precision. Our key finding—novel in the literature—is that users are highly sensitive to misrepresentations of critical concepts but consistently fail to recognize valid generalizations over non-critical concepts, frequently misclassifying them as model flaws; their explanation ratings show no significant discrimination between the two. This reveals a “abstraction-as-defect” cognitive bias in C-XAI interpretation, challenging prevailing faithfulness-centered evaluation paradigms and providing foundational cognitive insights for human-centered explainability design.
📝 Abstract
Concept-based explainable artificial intelligence (C-XAI) can help reveal the inner representations of AI models. Understanding these representations is particularly important in complex tasks like safety evaluation. Such tasks rely on high-level semantic information (e.g., about actions) to make decisions about abstract categories (e.g., whether a situation is dangerous). In this context, it may desirable for C-XAI concepts to show some variability, suggesting that the AI is capable of generalising beyond the concrete details of a situation. However, it is unclear whether people recognise and appreciate such generalisations and can distinguish them from other, less desirable forms of imprecision. This was investigated in an experimental railway safety scenario. Participants evaluated the performance of a simulated AI that evaluated whether traffic scenes involving people were dangerous. To explain these decisions, the AI provided concepts in the form of similar image snippets. These concepts differed in their match with the classified image, either regarding a highly relevant feature (i.e., relation to tracks) or a less relevant feature (i.e., actions). Contrary to the hypotheses, concepts that generalised over less relevant features led to ratings that were lower than for precisely matching concepts and comparable to concepts that systematically misrepresented these features. Conversely, participants were highly sensitive to imprecisions in relevant features. These findings cast doubts on whether people spontaneously recognise generalisations. Accordingly, they might not be able to infer from C-XAI concepts whether AI models have gained a deeper understanding of complex situations.