Comparing Imputation Methods for Clinical Prediction Model Development under Complex Missingness Scenarios: A Simulation Study Using Real-World Cardiac Data
This study addresses the challenge of diminished stability and generalizability of clinical prediction models under complex missing data, where the impact of different imputation strategies remains unclear. Leveraging a real-world cardiac disease cohort, we simulated 18 distinct missingness mechanisms to systematically evaluate how multiple imputation, missForest, k-nearest neighbors (kNN) imputation, and complete-case analysis affect logistic regression model performance. Model assessment encompassed internal and external validation metrics including AUC, calibration slope, prediction error, and computational efficiency. Our work provides the first comprehensive comparison of imputation methods across diverse missing data patterns, revealing that kNN imputation demonstrates superior robustness—particularly under high missingness rates and complex missingness structures—while achieving excellent external generalizability and the lowest computational cost, making it especially suitable for large-scale clinical modeling.