Position: Privacy Is a Claim, Not a Property of Synthetic Data

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了合成数据在隐私敏感环境中的使用问题,指出当前做法缺乏明确的威胁模型和可验证的隐私声明,并建议将隐私视为基于证据的科学声明。
📝 Abstract
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.
Problem

Research questions and friction points this paper is trying to address.

synthetic data
privacy
inference risk
threat models
privacy claims
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic data
privacy claim
threat models
inference risks
evidence-based
🔎 Similar Papers