A repeated k-fold cross-validation approach for evaluating the instability of clinical prediction models: an empirical comparison to the bootstrap approach

📅 2026-07-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic comparison between cross-validation and bootstrapping for assessing instability in clinical prediction models. Leveraging a cohort of 19,418 emergency department patients, it presents the first comprehensive evaluation of repeated five-fold cross-validation versus bootstrapping across varying events-per-variable (EPV) scenarios, using logistic regression and random forest models. Performance was assessed via AUC, calibration slope, large-scale calibration, and mean absolute prediction error (MAPE). Results indicate that when EPV ≥ 30, both methods yield comparable discriminative ability; however, cross-validation provides more accurate calibration estimates and significantly lower MAPE. These advantages render cross-validation particularly suitable for evaluating model instability across multiple algorithms, offering a dual benefit of internal validation and quantification of predictive stability.
📝 Abstract
Bootstrap-based methods have been recommended for assessing prediction instability in clinical prediction models, but their performance relative to cross-validation (CV) remains unclear. We propose a CV-based approach for assessing prediction instability and compare it with a bootstrap-based approach in logistic regression and random forest models. We conducted a resampling-based empirical experiment using a clinical cohort of 19,418 emergency department patients. Development samples were generated under events-per-variable (EPV) scenarios of 10, 30, and 50, and results were compared with those from the full dataset. Models were evaluated using bootstrap validation and repeated 5-fold CV; nested CV was used for random forest tuning. Predictive performance was assessed using AUC, calibration slope, and calibration-in-the-large. Prediction instability was quantified using mean absolute prediction error (MAPE). For logistic regression, bootstrap validation and repeated 5-fold CV produced broadly similar discrimination and calibration, especially at higher EPV values. For random forest, apparent performance consistently overestimated empirical discrimination. Bootstrap validation and repeated 5-fold CV gave comparable discrimination, but repeated 5-fold CV produced calibration slope estimates closer to the empirical value. Prediction stability improved as EPV increased for both modelling approaches. At EPV 30, bootstrap-derived MAPE was higher than CV-derived MAPE for both logistic regression (median, 0.042 versus 0.020) and random forest (median, 0.077 versus 0.027). A CV-based approach can assess prediction instability while also providing internally validated performance. These findings support CV-based instability assessment as a practical alternative to bootstrap-based assessment, particularly when comparing instability across multiple modelling algorithms.
Problem

Research questions and friction points this paper is trying to address.

prediction instability
clinical prediction models
cross-validation
bootstrap
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

repeated k-fold cross-validation
prediction instability
bootstrap validation
calibration slope
mean absolute prediction error
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Nop Khongthon
Department of Biomedical Informatics and Clinical Epidemiology (BioCE), Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
P
Pakpoom Wongyikul
Department of Biomedical Informatics and Clinical Epidemiology (BioCE), Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
N
Noraworn Jirattikanwong
Department of Biomedical Informatics and Clinical Epidemiology (BioCE), Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
P
Phanu Prasankittirach
Department of Biomedical Informatics and Clinical Epidemiology (BioCE), Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
N
Natthanaphop Isaradech
Department of Community Medicine, Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
W
Wachiranun Sirikul
Department of Community Medicine, Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.
N
Noppadon Seesuwan
Department of Emergency Medicine, Lampang Hospital, Lampang, Thailand.
S
Suppachai Lawanaskol
Chaiprakarn Hospital, Chiang Mai, Thailand.
P
Phichayut Phinyo
Department of Biomedical Informatics and Clinical Epidemiology (BioCE), Faculty of Medicine, Chiang Mai University, Chiang Mai, Thailand.