Institution profile

AryaXAI.com

Industry research
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Interpretability-Aware Pruning for Efficient Medical Image Analysis

Jul 11, 2025

Deep neural models for medical image analysis suffer from excessive parameter counts and poor interpretability, hindering clinical deployment. To address this, we propose an interpretability-aware structured pruning framework that—uniquely—leverages attribution methods (DL-Backprop, LRP, and Integrated Gradients) as dynamic pruning guidance signals. These signals identify and preserve neuron-level components critical for clinical diagnosis, enabling simultaneous model compression and joint optimization of predictive performance and decision transparency. Evaluated across multiple medical image classification benchmarks, our method achieves aggressive pruning (>50% parameter reduction) with negligible accuracy degradation (<0.5% drop), while substantially improving inference efficiency and interpretability. The framework establishes a novel paradigm for deploying lightweight, trustworthy AI models in clinical settings.

0 citationsRead paper

Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance

Feb 07, 2025

Current XAI evaluation lacks standardized, reliable, and manipulation-resistant metrics—primarily due to the absence of ground-truth explanations, resulting in low credibility, high regulatory compliance risks, and difficulties in cross-method comparison. To address this, we propose a novel evaluation paradigm that is human-centered, scenario-adaptive, and manipulation-resistant, integrating human-subject experiments, domain-knowledge modeling, regulatory alignment analysis, and principled benchmark design. We develop domain-specific evaluation benchmarks for high-stakes applications—including healthcare and finance—and advance a multi-stakeholder standardization framework. Our approach shifts XAI evaluation from subjective, fragmented practices toward a systematic, trustworthy, compliant, and comparable methodology. The framework enables scalable, empirically grounded assessment aligned with real-world deployment requirements and regulatory expectations. (149 words)

0 citationsRead paper

xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods

Feb 05, 2025

To address the lack of rigorous evaluation criteria for post-hoc explanation methods for black-box models, this paper introduces the first open-source, multimodal evaluation framework supporting both tabular and image data. The framework unifies three core dimensions—faithfulness, sensitivity, and robustness—into a reproducible, systematic benchmarking protocol. It integrates major explanation methods including SHAP, LIME, Grad-CAM, Integrated Gradients, and Backpropagation-based Trace (Backtrace), and supports models implemented in PyTorch and TensorFlow, as well as user-defined explainers. Extensive experiments on UCI tabular benchmarks and ImageNet demonstrate substantial performance variation across methods, underscoring the necessity of standardized evaluation. Our framework significantly improves the efficiency, comparability, and reproducibility of explanation credibility assessment. The implementation is publicly released and has been widely adopted by the research community.

0 citationsRead paper
Recent publications

Latest Papers

Interpretability-Aware Pruning for Efficient Medical Image Analysis

Jul 11, 2025

Deep neural models for medical image analysis suffer from excessive parameter counts and poor interpretability, hindering clinical deployment. To address this, we propose an interpretability-aware structured pruning framework that—uniquely—leverages attribution methods (DL-Backprop, LRP, and Integrated Gradients) as dynamic pruning guidance signals. These signals identify and preserve neuron-level components critical for clinical diagnosis, enabling simultaneous model compression and joint optimization of predictive performance and decision transparency. Evaluated across multiple medical image classification benchmarks, our method achieves aggressive pruning (>50% parameter reduction) with negligible accuracy degradation (<0.5% drop), while substantially improving inference efficiency and interpretability. The framework establishes a novel paradigm for deploying lightweight, trustworthy AI models in clinical settings.

0 citationsRead paper

Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance

Feb 07, 2025

Current XAI evaluation lacks standardized, reliable, and manipulation-resistant metrics—primarily due to the absence of ground-truth explanations, resulting in low credibility, high regulatory compliance risks, and difficulties in cross-method comparison. To address this, we propose a novel evaluation paradigm that is human-centered, scenario-adaptive, and manipulation-resistant, integrating human-subject experiments, domain-knowledge modeling, regulatory alignment analysis, and principled benchmark design. We develop domain-specific evaluation benchmarks for high-stakes applications—including healthcare and finance—and advance a multi-stakeholder standardization framework. Our approach shifts XAI evaluation from subjective, fragmented practices toward a systematic, trustworthy, compliant, and comparable methodology. The framework enables scalable, empirically grounded assessment aligned with real-world deployment requirements and regulatory expectations. (149 words)

0 citationsRead paper

xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods

Feb 05, 2025

To address the lack of rigorous evaluation criteria for post-hoc explanation methods for black-box models, this paper introduces the first open-source, multimodal evaluation framework supporting both tabular and image data. The framework unifies three core dimensions—faithfulness, sensitivity, and robustness—into a reproducible, systematic benchmarking protocol. It integrates major explanation methods including SHAP, LIME, Grad-CAM, Integrated Gradients, and Backpropagation-based Trace (Backtrace), and supports models implemented in PyTorch and TensorFlow, as well as user-defined explainers. Extensive experiments on UCI tabular benchmarks and ImageNet demonstrate substantial performance variation across methods, underscoring the necessity of standardized evaluation. Our framework significantly improves the efficiency, comparability, and reproducibility of explanation credibility assessment. The implementation is publicly released and has been widely adopted by the research community.

0 citationsRead paper