🤖 AI Summary
This study addresses the lack of purpose-specific compliance verification in evaluating synthetic tabular data by proposing a Locked Audit Framework. This approach fixes intended use and tolerance prior to evaluation, decoupling utility, privacy, and human authorization through pre-audit rules, variance adaptation, and anytime-valid certificates. It establishes query budget lower bounds to generate reproducible release evidence. Experiments across multiple datasets demonstrate that the framework effectively discriminates between model performance while tightening certificate bounds by two- to ten-fold. These results validate both the discriminative power and theoretical necessity of the proposed criteria, establishing a rigorous auditing paradigm for safe, context-aware data release in specific application scenarios.
📝 Abstract
Synthetic tabular data are often judged by realism, privacy, or downstream-task scores. Those scores do not answer whether a proposed release is supported for a named use, population, and threat model. We introduce SynthGuard-ReleaseBench, an audit framework that locks the use, candidate panel, tolerances, and audit schedule before evaluation. It compares real-trained and synthetic-trained workflows on protected data, gives simultaneous finite-sample bounds for bounded loss gaps, requires controls, and keeps utility, empirical privacy risk, mechanism claims, and human release authority separate.
Across four American Community Survey studies, five non-ACS records, two chronological diagnostics, and a sealed prototype, the benchmark retains favorable, unfavorable, and excluded outcomes. Transparent baselines pass some locked audits; compact learned models fail under the declared budgets; a health-table case is excluded because its negative control passes. A post-audit scaling arm, repeated across three generation seeds, shows the same locked criterion admitting those learned models once they are fit on enough data while still rejecting a dependence-destroying control at every size, so the criterion discriminates rather than merely rejects; the same repetition withdraws a finer single-seed ordering.
The theory adds a pre-audit sample-size rule, variance-adaptive and anytime-valid certificates that tighten the bound two to ten times on the same locked evidence, a temporal certificate for time-ordered audits, and two lower bounds: ordinary bounded queries reconstruct a protected audit once the query budget reaches its size, and the panel-size correction is necessary rather than conservative. The contribution is a reproducible workflow for use-specific release evidence, not a claim that any generator is private, safe, or deployment-ready.