Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing abductive explanation methods confined to Euclidean spaces, which hinders their adaptation to non-Euclidean prototypical networks. We extend the abductive latent explanation framework to non-Euclidean geometric manifolds by deriving geometric variant mappings and constructing specialized boundary algorithms to compute subset-minimal formal explanations. This work represents the first extension of formal explanations into non-Euclidean spaces, establishing a unified comparison framework for cross-architecture interpretability. Furthermore, we validate these theoretical constructions on image classification tasks, achieving the first rigorous comparative analysis of interpretability across different architectures. Collectively, this research bridges a critical gap in explainable AI by generalizing formal abductive reasoning beyond Euclidean constraints.
📝 Abstract
Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations - such as spherical metrics, Gaussian densities, and dimensional projections - rendering current formal explanation methods incompatible. In this work, we generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, we systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. We validate our theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, we enable the first rigorous, cross-architecture comparison of their interpretability.
Problem

Research questions and friction points this paper is trying to address.

Abductive Latent Explanations
Non-Euclidean representations
Prototype-based neural networks
Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Abductive Latent Explanations
Non-Euclidean Prototype Architectures
Formal Explanations
Interpretability
🔎 Similar Papers
No similar papers found.