🤖 AI Summary
To address low covariate efficiency and poor scalability to high-dimensional settings in regression discontinuity (RD) designs, this paper proposes a novel class of covariate-adjusted estimators. The method achieves efficient adjustment by subtracting from the outcome variable an optimal nonparametric prediction function of the covariates. Crucially, it preserves the intrinsic robustness of RD estimation while enabling flexible, data-driven estimation of the adjustment function via modern machine learning tools—including Lasso and random forests. Notably, it is the first RD adjustment framework that simultaneously attains asymptotic variance minimization and compatibility with machine learning estimators. Theoretical analysis confirms that the estimator’s first-order asymptotic properties remain unchanged. Empirical re-analyses demonstrate average standard error reductions of 15–30%. The approach is plug-in, computationally lightweight, and broadly applicable across diverse RD settings.
📝 Abstract
Empirical regression discontinuity (RD) studies often include covariates in their specifications to increase the precision of their estimates. In this paper, we propose a novel class of estimators that use such covariate information more efficiently than existing methods and can accommodate many covariates. Our estimators are simple to implement and involve running a standard RD analysis after subtracting a function of the covariates from the original outcome variable. We characterize the function of the covariates that minimizes the asymptotic variance of these estimators. We also show that the conventional RD framework gives rise to a special robustness property which implies that the optimal adjustment function can be estimated flexibly via modern machine learning techniques without affecting the first-order properties of the final RD estimator. We demonstrate our methods' scope for efficiency improvements by reanalyzing data from a large number of recently published empirical studies.