🤖 AI Summary
This paper addresses the bias induced by machine learning (ML) plug-in estimators in U-statistics—particularly inequality of opportunity (IOp)—and proposes the first debiased IOp estimator with accompanying asymptotic inference theory. Methodologically, it integrates double robustness, U-statistic theory, and general-purpose ML prediction (e.g., tree-based models, regression). The approach effectively corrects estimation bias arising from model selection and regularization. Contributions include: (1) a novel debiasing framework for general U-statistics, overcoming the sensitivity of conventional plug-in methods to ML model specification; (2) the first unbiased cross-national IOp estimates for European countries; and (3) empirical identification of maternal education and paternal occupation as the most critical circumstances shaping IOp. Simulation studies demonstrate substantial improvements in estimation accuracy and reliable statistical inference.
📝 Abstract
Equality of opportunity has emerged as an important ideal of distributive justice. Empirically, Inequality of Opportunity (IOp) is measured in two steps: first, an outcome (e.g., income) is predicted given individual circumstances; and second, an inequality index (e.g., Gini) of the predictions is computed. Machine Learning (ML) methods are tremendously useful in the first step. However, they can cause sizable biases in IOp since the bias-variance trade-off allows the bias to creep in the second step. We propose a simple debiased IOp estimator robust to such ML biases and provide the first valid inferential theory for IOp. We demonstrate improved performance in simulations and report the first unbiased measures of income IOp in Europe. Mother's education and father's occupation are the circumstances that explain the most. Plug-in estimators are very sensitive to the ML algorithm, while debiased IOp estimators are robust. These results are extended to a general U-statistics setting.