Subgroup Membership Inference Audits of Differentially Private Synthetic Text

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过定义子组目标成员推理游戏,审计差分隐私合成文本对子组成员信息的泄露风险,发现现有方法低估了该风险。
📝 Abstract
Synthetic data releases are increasingly proposed in the literature as a means of sharing realistic data replicas in lieu of sensitive private datasets. Even when the worst-case privacy leakage of such releases is bounded by means of differential privacy (DP), in practice a residual risk remains. Membership inference attack (MIA) audits are conducted to empirically quantify this risk. However, existing methods only measure average-case risk for randomly drawn records, which might conceal the risk to vulnerable subgroups. To highlight this issue, we define a subgroup-targeted membership inference game in which the target pool is an explicit parameter, and instantiate it with an audit of 32 proxies under three scenarios with different levels of attacker knowledge, across four datasets, three generators (DP-SGD fine-tuning, API-based prompting, and activation steering), and five privacy budgets. The audit shows that synthetic releases leak subgroup membership and that prior attacks systematically underestimate this leakage. DP is effective at the aggregate level: it substantially reduces average leakage at every budget we test. Three observations temper this picture. First, the remaining leakage is concentrated rather than spread out: under DP, a tenth of the records carries roughly 40% of it. Second, the protection DP delivers in practice is uneven: within its worst-case guarantee, the noise removes more of the measured leakage from random records than from high-risk ones---and a merged-pool audit that scores both record types against shared negatives confirms this at the record level. Third, \emph{which} records leak proves to be a property of the release mechanism rather than of the record alone, so record-level risk cannot be assessed independently of the release.
Problem

Research questions and friction points this paper is trying to address.

Membership Inference Attack
Differential Privacy
Synthetic Data
Subgroup Risk
Privacy Leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

subgroup-targeted membership inference
differential privacy
synthetic data releases
privacy leakage
🔎 Similar Papers
Y
Yidan Sun
Imperial College London, Imperial Global Singapore
Viktor Schlegel
Viktor Schlegel
Deputy Director IN-CYPHER Programme @ IGS, Imperial College London
Natural Language UnderstandingAI for HealthcareClinical NLPAI Evaluation
S
Srinivasan Nandakumar
Imperial College London, Imperial Global Singapore
S
Siew Kei Lam
Nanyang Technological University, Singapore
A
Anil Anthony Bharath
Imperial College London, United Kingdom