Recent Advances in Medical Image Classification
Medical image classification faces two critical challenges: scarcity of labeled samples in clinical settings and insufficient interpretability for medical practitioners. This paper presents a systematic review of recent advances and proposes a novel few-shot learning framework that synergistically integrates Vision Transformers (ViTs) with vision-language models (VLMs)—the first such integration for medical image classification. To enhance clinical trustworthiness, the framework incorporates explainable AI (XAI) techniques to generate human-interpretable visualizations (e.g., attention heatmaps) and natural-language explanations aligned with clinical reasoning. Extensive experiments on multiple public medical imaging benchmarks demonstrate that our method significantly improves few-shot classification accuracy—achieving an average gain of +5.2%—while delivering clinically coherent, multimodal explanations. The approach advances the practical deployment of AI-assisted diagnosis by bridging the gap between high-performance deep learning and domain-specific interpretability requirements.