🤖 AI Summary
This paper addresses the unreliability of causal and predictive parameter estimation under covariate shift. We propose a fully automated debiasing machine learning framework that eliminates regularization bias solely through parameter definition—without requiring explicit bias modeling. Our approach innovatively integrates training and target data within a unified debiasing mechanism, combining data fusion, high-dimensional statistical inference, doubly robust estimation, and the difference-in-differences (DID) principle—all under an unconfoundedness assumption. We establish theoretical guarantees of consistency and asymptotic normality. In simulation studies and an empirical analysis of minimum wage effects on teenage employment, our method reduces estimation bias by over 40% on average compared to benchmark approaches, while substantially improving estimation accuracy and robustness.
📝 Abstract
We present machine learning estimators for causal and predictive parameters under covariate shift, where covariate distributions differ between training and target populations. One such parameter is the average effect of a policy that alters the covariate distribution, such as a treatment modifying surrogate covariates used to predict long-term outcomes. Another example is the average treatment effect for a population with a shifted covariate distribution, like the effect of a policy on the treated group. We propose a debiased machine learning method to estimate a broad class of these parameters in a statistically reliable and automatic manner. Our method eliminates regularization biases arising from the use of machine learning tools in high-dimensional settings, relying solely on the parameter's defining formula. It employs data fusion by combining samples from target and training data to eliminate biases. We prove that our estimator is consistent and asymptotically normal. Computational experiments and an empirical study on the impact of minimum wage increases on teen employment--using the difference-in-differences framework with unconfoundedness--demonstrate the effectiveness of our method.