EEG-PRISM: Physiologically-Grounded Interpretability of Predictions by EEG Foundation Models

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of clinically intuitive, physiologically grounded attribution in EEG foundation models by proposing a universal post-hoc attribution framework that requires no retraining. Leveraging linear transformations, backpropagation, and invertible mappings, the method projects model predictions onto frequency and source domains to enable transient event localization and biomarker identification. Experiments demonstrate near-perfect recovery of simulated spectra with 69.2% spatial accuracy. Furthermore, the framework achieves 50% accuracy in localizing epileptic seizure onset zones and identifies autism biomarkers consistent with prior literature. These results significantly enhance the clinical applicability of model interpretability by bridging the gap between complex deep learning representations and established neurophysiological concepts, thereby facilitating more trustworthy deployment of EEG foundation models in clinical diagnostics and neuroscience research.
📝 Abstract
Objective: Foundation models represent the next advancement in AI for EEG analysis; however current explainable AI techniques provide attribution scores in the time-channel input space, which is mismatched to clinical intuition about EEG. Thus, there is a critical need for a universal method that can extend the interpretability of any foundation model to alternative and physiologically relevant domains without modifying or retraining the underlying model. Methods: EEG-PRISM leverages linear transformations and established backpropagation rules to map time-channel attribution scores into alternative domains. We derive mappings to the frequency domain via an invertible DFT and to the source domain via an approximately invertible EEG generative model. We evaluate EEG-PRISM in simulated and real data, assessing recovery of ground-truth phenomena across domains with five foundation models and four AI explainers. Results: In simulation, EEG-PRISM achieves near-perfect spectral recovery and 69.2% spatial accuracy. In epilepsy, EEG-PRISM correctly determines that delta-theta activity is most salient and correctly localizes the seizure onset region with 50% accuracy. In autism, EEG-PRISM localizes the predictive delta-alpha biomarkers to frontal and temporal regions, consistent with prior work. Conclusion: EEG-PRISM is a theoretically-grounded post-hoc attribution method with accurate mapping into the spectral and spatial domains. It supports window-level analysis of transient events (e.g., seizures) and group-level identification of clinically relevant biomarkers (e.g., autism), thus advancing interpretable EEG foundation models. Significance: This work enables physiologically-grounded interpretation of EEG foundation models and supports clinically relevant insights such as event localization and biomarker identification.
Problem

Research questions and friction points this paper is trying to address.

EEG foundation models
interpretability
physiologically-grounded
explainable AI
attribution mapping
Innovation

Methods, ideas, or system contributions that make the work stand out.

EEG Foundation Models
Physiologically-Grounded Interpretability
Post-hoc Attribution
Domain Mapping
Clinical Biomarkers
D
Deeksha M. Shama
Department of Electrical and Computer Engineering, Johns Hopkins University, MD, USA 21218; and visiting student researcher at Boston University, MA, USA 02215
P
Punnisa Amornsirikul
Boston University, MA, USA 02215
Archana Venkataraman
Archana Venkataraman
Boston University
Medical Image AnalysisComputational NeuroscienceMachine Learning