Bias-corrected Cox regression with AI-extracted covariates via calibration summary statistics
This study addresses the bias in Cox proportional hazards model estimates arising from measurement error in AI-extracted covariates, a setting where downstream users only have access to the extracted data and limited calibration summary statistics. Within a multivariate calibration framework, the work provides the first decomposition of Cox model bias into a dominant, calibratable component and higher-order residual terms. Building on this insight, the authors propose a post-processing correction method that relies solely on calibration summary statistics and can be directly applied to outputs from standard Cox regression software. The approach is accompanied by uncertainty-adjusted confidence intervals and sensitivity diagnostic tools. Empirical evaluations on synthetic data demonstrate substantial bias reduction, with near-nominal coverage maintained even under mild violations of the linear calibration assumption. The paper also recommends a minimal set of calibration statistics that data providers should report to enable effective bias correction.