Learning a Size-Weight Frontier for Synthetic-Augmented Inference

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了合成数据增强推理中的偏差问题,通过构建一个大小-权重前沿框架来控制合成样本数量和权重,从而确保目标覆盖率并缩小置信区间。
📝 Abstract
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.
Problem

Research questions and friction points this paper is trying to address.

synthetic data
statistical inference
bias
unreliable inference
augmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic-augmented inference
size-weight frontier
task-marginal coverage
finite-sample coverage guarantee
🔎 Similar Papers