Inferential Evaluation of Surrogate-Derived Models under Covariate Shift

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of surrogate model evaluation under covariate shift when target domain labels are unavailable. We propose a three-sample cross-fitting estimator based on density ratio estimation and kernel calibration. By integrating outcome regression augmentation, this method establishes asymptotic linear inference theory for TPR, FPR, and AUC, achieving consistent ROC curve inference and AUC asymptotic normality. This approach effectively resolves performance evaluation bottlenecks in transfer learning. Validated on real-world AI applications such as Chatbot Arena, the proposed framework provides a rigorous statistical inference foundation for reliable model assessment in unsupervised domain adaptation scenarios.
📝 Abstract
In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved. Evaluating its target performance is essential for determining whether decisions based on the model remain reliable, yet it is difficult when gold labels are scarce, and covariate distributions differ across data sources. We study a three-sample setting with a small gold-labeled source, a larger surrogate-labeled source, and an unlabeled target. Under conditional transportability, we evaluate the surrogate-derived model against the latent gold-standard outcome in the target population. We propose cross-fitted estimators that transport information from the two labeled sources through source-specific density ratios. We also combine outcome-regression augmentation with a kernel correction for estimating the model near a threshold, accounting for uncertainty from all three samples. We establish asymptotically linear inference for TPR and FPR, consistency and pointwise inference for the ROC curve, and asymptotically normal inference for AUC. Simulations assess bias, coverage, and sensitivity to bandwidth and relative sample sizes. A retrospective temporal validation on Chatbot Arena and a semi-synthetic ACS-Income study provide validation in real-world AI applications.
Problem

Research questions and friction points this paper is trying to address.

Covariate Shift
Surrogate Labels
Model Evaluation
Transfer Learning
Gold-standard Outcomes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Covariate Shift
Cross-fitted Estimators
Surrogate Labels
Asymptotic Inference
Transfer Learning Evaluation