Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the clinical challenges of confidence miscalibration and unreliable predictions in skin lesion classification by constructing a shared-backbone dual-head multi-task architecture trained on heterogeneous datasets. We systematically evaluated five uncertainty quantification methods, revealing that hard sample identification exhibits significant method-agnostic properties and that deferring high-uncertainty cases effectively reduces misdiagnosis rates. Notably, Deep Ensembles demonstrated superior performance in both model calibration and uncertainty decomposition. These findings validate the efficacy of uncertainty quantification in supporting selective referral strategies within clinical workflows. Ultimately, this work provides empirical evidence for enhancing the safety and reliability of AI-assisted diagnostic systems by integrating robust uncertainty estimation to mitigate risks associated with overconfident erroneous predictions in dermatological applications.
📝 Abstract
Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction. We study uncertainty quantification on a dataset pooled from many ISIC sources, with a shared backbone and two jointly learned heads: a binary malignant versus non-malignant head and a five-class diagnostic head. Five UQ methods (MC Dropout, DropConnect, Flipout, Deep Ensembles, DUQ) are compared on accuracy, calibration, uncertainty decomposition, and risk-coverage. Difficulty is largely method-agnostic: even methods with narrow entropy distributions rank the same samples as hard (per-sample entropy correlations of $0.54$ to $0.91$). The choice of method matters more for calibration and uncertainty decomposition, where Deep Ensembles is the clear winner, than for finding difficult cases. The ranking is also good enough that deferring the most uncertain cases removes a disproportionate share of errors, supporting uncertainty-based selective referral, evaluated here in-distribution only.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Quantification
Skin Lesion Classification
Difficult Sample Identification
Selective Referral
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Quantification
Method-Agnostic Difficulty
Multi-Task Learning
Selective Referral
Deep Ensembles
🔎 Similar Papers
No similar papers found.