GastroViT: A Vision Transformer Based Ensemble Learning Approach for Gastrointestinal Disease Classification with Grad CAM & SHAP Visualization

📅 2025-09-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of fine-grained classification of gastrointestinal endoscopic images by proposing GastroViT, a lightweight and interpretable ensemble model. Methodologically, it integrates two pretrained vision transformer backbones—MobileViT_XS and MobileViT_V2_200—and enhances feature discriminability via attention mechanisms. Dual-path interpretability is achieved through synergistic Grad-CAM (for lesion localization) and SHAP (for quantitative feature attribution). On 23-class and 16-class gastrointestinal disease classification tasks, GastroViT achieves 91.98% (F1 = 64%) and 92.70% (F1 = 87%) accuracy, respectively, with only 20 million parameters, no data augmentation, and robustness to class imbalance. Its core contribution lies in being the first to jointly leverage Vision Transformer ensembling and multimodal attribution-based visualization—thereby unifying high accuracy, low computational complexity, and clinical interpretability—establishing a novel paradigm for AI-assisted early screening in gastroenterology.

Technology Category

Application Category

📝 Abstract
The gastrointestinal (GI) tract of humans can have a wide variety of aberrant mucosal abnormality findings, ranging from mild irritations to extremely fatal illnesses. Prompt identification of gastrointestinal disorders greatly contributes to arresting the progression of the illness and improving therapeutic outcomes. This paper presents an ensemble of pre-trained vision transformers (ViTs) for accurately classifying endoscopic images of the GI tract to categorize gastrointestinal problems and illnesses. ViTs, attention-based neural networks, have revolutionized image recognition by leveraging the transformative power of the transformer architecture, achieving state-of-the-art (SOTA) performance across various visual tasks. The proposed model was evaluated on the publicly available HyperKvasir dataset with 10,662 images of 23 different GI diseases for the purpose of identifying GI tract diseases. An ensemble method is proposed utilizing the predictions of two pre-trained models, MobileViT_XS and MobileViT_V2_200, which achieved accuracies of 90.57% and 90.48%, respectively. All the individual models are outperformed by the ensemble model, GastroViT, with an average precision, recall, F1 score, and accuracy of 69%, 63%, 64%, and 91.98%, respectively, in the first testing that involves 23 classes. The model comprises only 20 million (M) parameters, even without data augmentation and despite the highly imbalanced dataset. For the second testing with 16 classes, the scores are even higher, with average precision, recall, F1 score, and accuracy of 87%, 86%, 87%, and 92.70%, respectively. Additionally, the incorporation of explainable AI (XAI) methods such as Grad-CAM (Gradient Weighted Class Activation Mapping) and SHAP (Shapley Additive Explanations) enhances model interpretability, providing valuable insights for reliable GI diagnosis in real-world settings.
Problem

Research questions and friction points this paper is trying to address.

Classifying gastrointestinal diseases from endoscopic images
Developing ensemble vision transformers for medical diagnosis
Enhancing interpretability with explainable AI visualization methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ensembles pre-trained vision transformers for disease classification
Combines MobileViT models with 20M parameters efficiently
Integrates Grad-CAM and SHAP for explainable AI visualization
S
Sumaiya Tabassum
Rajshahi University of Engineering & Technology, Rajshahi -6204, Bangladesh
M
Md. Faysal Ahamed
Rajshahi University of Engineering & Technology, Rajshahi -6204, Bangladesh
Hafsa Binte Kibria
Hafsa Binte Kibria
Assistant Professor, Department of Electrical & Computer Engineering, RUET
Machine LearningData MiningAI
M
Md. Nahiduzzaman
Rajshahi University of Engineering & Technology, Rajshahi -6204, Bangladesh
J
Julfikar Haider
Department of Engineering, Manchester Metropolitan University, Chester Street, Manchester M1 5GD, UK
M
Muhammad E. H. Chowdhury
Qatar University
M
Mohammad Tariqul Islam
Universiti Kebangsaan Malaysia