🤖 AI Summary
This paper addresses the estimation of the average treatment effect (ATE) under partial treatment-status missingness—where only the treated units are labeled, and all others constitute unlabeled data—a problem at the intersection of causal inference and positive-unlabeled (PU) weakly supervised learning. We first derive the semiparametric efficient bound for PU-type ATE estimation and construct a doubly robust estimator achieving this bound. Our estimator integrates influence function theory, regularized machine learning, and the PU learning framework, ensuring $sqrt{n}$-consistency and asymptotic normality. Theoretically, it attains the minimal asymptotic variance among all regular estimators. Empirical evaluations—including simulations and real-data analysis—demonstrate its substantial improvement over existing approaches that either discard unlabeled samples or rely on misspecified models. To our knowledge, this is the first method provably achieving semiparametric efficiency for ATE estimation under PU sampling, thereby providing the first theoretically optimal solution for weakly supervised causal inference.
📝 Abstract
The estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficient estimators for ATE estimation in a setting where only a treatment group and an unknown group-comprising units for which it is unclear whether they received the treatment or control-are observable. This scenario represents a variant of learning from positive and unlabeled data (PU learning) and can be regarded as a special case of ATE estimation with missing data. For this setting, we derive semiparametric efficiency bounds, which provide lower bounds on the asymptotic variance of regular estimators. We then propose semiparametric efficient ATE estimators whose asymptotic variance aligns with these efficiency bounds. Our findings contribute to causal inference with missing data and weakly supervised learning.