🤖 AI Summary
该研究提出了一种基于恒定步长随机梯度下降的检测器,用于识别潜在混合模型中的子群体,以避免因混合不同子群体而导致的误导性回归结论。
📝 Abstract
Pooling latent subpopulations can obscure relationships and yield misleading regression conclusions, including Simpson's paradox (SP). We propose a detector based on the steady-state dynamics of constant-step stochastic gradient descent (SGD). Unlike likelihood-based mixture tests and confounder-search methods, it requires neither a normal-mixture specification nor observed candidate confounders. Two linear pieces compete under a winner-take-all squared-error loss, and their normalized terminal separation forms the test statistic. Using diffusion approximations, we derive its asymptotic null distribution for general centered scalar covariates and for Gaussian multivariate covariates under symmetric noise. The distribution-dependent null center reduces to the dimension-free constant $4/\pi$ when both covariates and noise are Gaussian. We establish asymptotic size control and consistency against fixed mixture alternatives under regularity conditions. We extend the method to intercept and partial-mixture heterogeneity and study endogeneity, heteroskedasticity, and nonlinear misspecification. Simulations examine calibration, power, and robustness. Finally, a three-stage Detect--Screen--Verify toolkit separates evidence of heterogeneity from its substantive explanation. Across eight public datasets, it recovers four established SP benchmarks and identifies four cases that, to our knowledge, have not been documented previously. The detector requires neither latent-group labels nor a prespecified number of mixture components.