🤖 AI Summary
本文提出一种基于条件分布的框架,用于验证合成多变量数据是否保持了目标群体的联合依赖结构,并通过多个案例评估了该方法的有效性。
📝 Abstract
Statistical validation of synthetic multivariate data requires assessing whether a generator preserves the joint dependence structure of the target population without merely reproducing observed records. We develop a model-agnostic framework based on full conditional distributions. For each coordinate, we normalize the conditional probability assigned to the observed value by the largest conditional probability available in the same record context; averaging this quantity yields a one-sided MAP-alignment statistic that can be estimated using a conditional model fitted on held-out real data. The mathematical contribution is twofold: under strict positivity and compatibility, the complete normalized conditional profile identifies the joint distribution, and its integrated L1 difference defines a metric on finite-state generative processes; we also establish consistency and finite-sample concentration for the corresponding empirical estimators. Because high conditional alignment alone can arise from copying or concentration on conditional modes, we pair it with nearest-real similarity as a separate record-level novelty diagnostic. We evaluate the framework on NSHAP health and aging data, influenza B genomic surveillance, and 34 General Social Survey waves. In GSS, the Large Science Model matched the original-data control in mean conditional alignment while retaining substantial novelty, indicating preservation of conditional structure without row reuse. In influenza B, a Chow-Liu generator matched the control alignment but had almost no novelty, revealing near-reproduction of observed records. The framework therefore distinguishes three statistically different failure modes: loss of dependence, record reuse, and mode concentration, and provides a principled basis for validating synthetic health, surveillance, and population data.