ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
ProxyGuard通过预设风险和封闭目标集控制错误,评估共享目标机制的可靠性,提高研究中随机数据发布机制的有效性和可靠性。
📝 Abstract
Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism draws on a common target, lower-bounds their favorable-score rate, and subtracts a bound on favorable scores contributed by invalid releases. Conditional on the target, release scores are independent, yielding a finite-sample mechanism-reliability guarantee without independent target batches or assumptions on release-level $p$-value dependence. We show that the mean-only penalty is sharp and derive a smooth-score certificate with additive target concentration. In a registered three-requirement study, direct mode raises power from 5.6\% to 64.2\% at reliability 0.95, while named mode remains stronger under high-signal evidence. Prospective audits span full-pipeline Rice--TVAE, which retrains on every draw, and a non-tabular text mechanism.
Problem

Research questions and friction points this paper is trying to address.

ProxyGuard
reliability inference
randomized data release mechanisms
shared targets
multiplicity correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

ProxyGuard
reliability inference
randomized data release mechanisms
shared targets
finite-sample guarantee
🔎 Similar Papers
No similar papers found.
D
Dipesh Tharu Mahato
New York University
P
Pramod Dhungana
Queens College