UD-DML: Uniform Design Subsampling for Double Machine Learning over Massive Data

📅 2026-05-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of double machine learning (DML) with large-scale observational data and the inadequacy of simple random subsampling, which often results in insufficient covariate coverage and imbalance between treatment and control groups. The authors propose UD-DML, a novel subsampling approach that incorporates uniform design principles into DML for the first time. The method first applies PCA rotation to covariates, then constructs a low-discrepancy skeleton based on a hybrid discrepancy criterion, and finally uses KD-trees to match nearest-neighbor treated and control units to each skeleton point, yielding a representative and balanced subsample. Theoretical analysis establishes that UD-DML retains √n-asymptotic normality even when the subsample size is much smaller than the full dataset. Experiments demonstrate that, compared to uniform sampling, UD-DML achieves lower RMSE, narrower confidence intervals, and more reliable coverage—particularly under low overlap and model misspecification.
📝 Abstract
Double machine learning (DML) delivers valid inference on low-dimensional causal parameters while permitting flexible nuisance estimation, but its computational cost becomes prohibitive once cross-fitted learners must be trained on massive observational data. Applying DML to a uniformly drawn subsample alleviates this burden, yet such a reduction disregards the geometry of the covariate space and can exacerbate treated-control imbalance as well as overlap deficiency. We propose Uniform Design Double Machine Learning (UD-DML), a design-based subsampling strategy for average treatment effect (ATE) estimation. UD-DML first constructs a low-discrepancy skeleton in a PCA-rotated covariate space under the mixture-discrepancy criterion, and then assigns, to each skeleton point, the nearest treated and control units via KD-tree search. The resulting matched subsample is, by construction, both representative of the full covariate distribution and balanced across treatment arms; cross-fitted DML is subsequently applied to it. We establish discrepancy-based guarantees for representativeness and balance, and prove that the UD-DML estimator is $\sqrt{r}$-asymptotically normal under mild conditions, where the selected subsample size $r \ll n$. The dominant nuisance-fitting cost is thereby reduced from the $n$-scale to the $r$-scale. Monte Carlo experiments show that UD-DML attains lower RMSE, narrower confidence intervals and more reliable coverage than uniform subsampling, with the largest gains in low-overlap and misspecified regimes. An application to a large observational dataset further demonstrates its practical feasibility.
Problem

Research questions and friction points this paper is trying to address.

Double Machine Learning
Subsampling
Treatment Effect Estimation
Covariate Balance
Massive Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Double Machine Learning
Uniform Design
Subsampling
Low-discrepancy Sampling
Average Treatment Effect
🔎 Similar Papers
Y
Yuanke Qu
School of Computer Science and Engineering, Guangdong Ocean University, Guangdong 529500, China
X
Xiaoya Xu
Institute of Applied Mathematics, Shenzhen Polytechnic University, Shenzhen 518055, China
H
Hengtao Zhang
School of Computer Science and Engineering, Guangdong Ocean University, Guangdong 529500, China