๐ค AI Summary
This work addresses the critical gap in principled quantification of prediction uncertainty in clinical AI systems, which undermines their trustworthiness and fairness in high-stakes medical settings. The authors propose the first end-to-end Bayesian deep learning framework that integrates multimodal patient data through modality-specific variational encoders, precision-weighted late fusion, a composite Bayesian loss, and an uncertainty calibration penalty to disentangle aleatoric and epistemic uncertainties. Notably, they introduce calibrated uncertainty as a formal metric for algorithmic fairness and conduct cross-subgroup audits across demographic and socioeconomic strata. Experimental results demonstrate a low expected calibration error (ECE = 0.096) and reveal significant fairness gaps in uncertainty estimates for patients from primary/rural care settings, those with low socioeconomic status, and older adults (p < 0.001), while no significant disparity was observed across gender groups.
๐ Abstract
Clinical artificial intelligence (AI) systems routinely produce predictions without principled quantification of uncertainty, limiting their trustworthiness in high-stakes medical environments. This paper presents an integrated research programme addressing two interconnected problems: (1) the development of a fully end-to-end Bayesian uncertainty modelling framework for multimodal clinical data, and (2) the application of calibrated uncertainty estimates as a formal measure of algorithmic equity across patient subgroups. We construct a probabilistic deep learning architecture comprising modality-specific variational encoders, a precision-weighted late fusion mechanism, and a decomposed uncertainty output head that separates aleatoric from epistemic uncertainty. The system is trained with a composite Bayesian loss incorporating binary cross-entropy, Kullback-Leibler divergence regularisation, and an uncertainty calibration penalty. We evaluate model calibration using Expected Calibration Error (ECE = 0.096) and conduct a subgroup equity audit across facility type, socioeconomic status, age group, and biological sex on a dataset of 1,000 simulated patients. Results demonstrate that epistemic uncertainty systematically identifies underserved populations: primary/rural facility patients show a 15.3% uncertainty equity gap (p < 0.001, effect size = 0.698), low socioeconomic status patients exhibit a 6.8% gap (p < 0.001), and elderly patients show a 3.9% gap (p < 0.001), whilst no significant sex-based disparity is detected. These findings establish that calibrated uncertainty is not merely a technical property of probabilistic models but constitutes an actionable equity signal with direct clinical relevance.