🤖 AI Summary
该研究通过将训练子样本视为第二阶段抽样,解决了模型辅助估计中的不确定性量化问题,并提出了两种方差估计方法。
📝 Abstract
When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the training subsample as a second phase of sampling and derive, exactly and for any algorithm, a two-term variance decomposition and the variance family linking the single-partition estimator to its Rao-Blackwellized average, whose design bias it shares. For tree-type predictors the second-phase variance is computable in closed form, and its share of total variance grows with tree complexity, explaining documented variance underestimation. We propose an analytic and a replication variance estimator, neither altering the point estimate, and evaluate them by simulation: budgeting the second phase restores near-nominal coverage at a small fraction of the cost of partition averaging.