Cascade Classification of Dermoscopic Images of Skin Neoplasms with Controllable Sensitivity and External Clinical Validation

๐Ÿ“… 2026-06-11
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limited generalization and miscalibration of dermoscopic image classification models when transferring from international public datasets (e.g., ISIC Archive) to real-world clinical settings (e.g., Russian clinical data). To overcome the constraints of conventional single-stage argmax classification, the authors propose a tunable-sensitivity, two-stage cascaded framework: an initial benignโ€“malignant binary screening followed by fine-grained subclassification of malignant cases, with adjustable decision thresholds aligned to clinical diagnostic logic. Evaluated consistently across ViT-B/16, Swin-S, ConvNeXt-S, and EfficientNetV2-S architectures, the approach achieves internal test ROC-AUC scores of 0.952โ€“0.966 for binary classification. Although performance declines on external clinical data, the cascaded design significantly improves the macro F1 score for ViT-B/16 and effectively reduces missed diagnoses of malignant cases.
๐Ÿ“ Abstract
Purpose. To compare deep learning architectures and classification schemes for dermoscopic images of skin neoplasms and assess their generalization on transfer from open international datasets to independent clinical datasets of Russian practice. Methods. Four architectures (ViT-B/16, Swin-S, ConvNeXt-S, EfficientNetV2-S) were compared in three schemes: binary (malignant/benign), single-stage four-class (benign, MEL, SCC, BCC), and a two-stage cascade (binary triage, then three-class differentiation MEL/SCC/BCC). All models used ImageNet-pretrained weights and a single augmentation protocol on aggregated open ISIC Archive data, and were evaluated on an internal held-out sample and two clinical datasets (Melanoscope AI mobile system; Sechenov University). Results. Internally the binary stage attains ROC-AUC 0.952-0.966; on Sechenov University it drops to 0.797-0.893, sensitivity to 0.53-0.67, and ECE rises from 0.02 to 0.27-0.39 with underestimation of malignancy, quantifying a generalization gap in ranking and calibration. Paired tests confirm one inter-architecture result on clinical data: the deficit of ViT-B/16 at the binary stage (p<0.05); at the differentiation stage no architecture has a proven advantage. The cascade raises macro F1 over single-stage four-class classification for most architectures, but significantly only for ViT-B/16, by recovering malignant lesions assigned to the dominant benign class. On ISIC MILK10k, direct 11-class classification yields mean-class sensitivity 0.525. Conclusion. A tunable triage threshold gives sensitivity control not attainable in standard single-stage (argmax) classification and better reproduces clinical differential-diagnosis logic. The persistent generalization gap mandates external clinical validation and recalibration before deployment.
Problem

Research questions and friction points this paper is trying to address.

dermoscopic images
skin neoplasms
generalization gap
clinical validation
sensitivity control
Innovation

Methods, ideas, or system contributions that make the work stand out.

cascade classification
controllable sensitivity
external clinical validation
dermoscopic image analysis
generalization gap
๐Ÿ’ผ Related Jobs
No related jobs found.
E
Elena S. Kozachok
Ivannikov Institute for System Programming of the Russian Academy of Sciences (ISP RAS), Moscow, Russia
S
Sergey S. Seregin
Orel Oncological Dispensary, Orel, Russia
A
Aleksandr V. Kozachok
Ivannikov Institute for System Programming of the Russian Academy of Sciences (ISP RAS), Moscow, Russia
I
Ilya P. Latyshev
Ivannikov Institute for System Programming of the Russian Academy of Sciences (ISP RAS), Moscow, Russia
O
Oleg I. Samovarov
Ivannikov Institute for System Programming of the Russian Academy of Sciences (ISP RAS), Moscow, Russia