Significance-First Splitting: Aligning Treatment Heterogeneity Detection with Honest Estimation
Existing methods for heterogeneous treatment effect estimation often struggle to simultaneously ensure sensitivity in detecting effect modifiers and validity in statistical inference. This work proposes a hybrid algorithm that integrates significance-driven splitting with honest estimation: it employs the t² statistic as the splitting criterion, incorporates honest sample splitting, selects the cost-complexity penalty via cross-validation, and uses the infinitesimal jackknife to estimate Monte Carlo variance. This approach is the first to align significance-based splitting with an honest estimation framework, maintaining theoretical consistency under strong interactions and providing nominal leaf-level confidence intervals for a single tree. Empirical results demonstrate approximately 90% coverage (nominal 90%) on Athey–Imbens synthetic data and Qini coefficients on par with S- and T-learners on real-world Criteo and Starbucks datasets.