🤖 AI Summary
This study addresses the unclear statistical role and unverified spatial adaptivity of stopping rules in CART regression trees by integrating minimax rate analysis with nonparametric regression theory. The primary contribution lies in establishing, for the first time, the precise statistical function of the Minimum Impurity Decrease (MID) rule. We prove that MID achieves optimal pointwise convergence rates within logarithmic factors and global uniform spatial adaptivity, whereas the minimum leaf size rule fails to attain such performance. Consequently, this work provides a critical theoretical foundation for the CART algorithm and elucidates the fundamental differences between distinct stopping criteria in adaptive estimation.
📝 Abstract
The popular CART algorithm for regression trees combines a greedy splitting rule with a stopping rule, but while the splitting rule has been well studied, the statistical role of stopping rules is less well understood. Meanwhile, although regression trees fit using Bayesian methods or via empirical risk minimization (ERM) have been shown to be spatially adaptive to local smoothness and anisotropy, it is unknown whether CART can achieve the same adaptation. We address these gaps by proving that, under spatially heterogeneous and anisotropic smoothness and appropriate structural assumptions on the regression function and covariate distribution, CART with the minimum impurity decrease (MID) stopping rule and a suitable threshold achieves pointwise rates that are minimax up to logarithmic factors. These rates hold simultaneously over all points in the domain. Moreover, we prove that spatial adaptation cannot be achieved under the widely used minimum leaf size stopping rule. Together, these results establish a precise statistical role for the MID stopping rule and provide a theoretical basis for the empirical success of CART.