The Privacy-Hallucination Tradeoff in Differentially Private Language Models

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示了差分隐私语言模型中的隐私-幻觉权衡问题,通过实验分析其原因,并提出需更细致的隐私保护方法以保证事实准确性。
📝 Abstract
Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tuned with DP tend to produce more hallucinations than non-DP counterparts, with increased severity as the privacy budget grows stricter. Second, we investigate model properties driving this tradeoff, demonstrating that DP mechanisms flatten output distributions, potentially redistributing probability mass toward factually incorrect alternatives. Third, through experiments where we control fact frequency in training data, we characterize how information frequency can reduce hallucination risks in DP models. Overall, our findings underscore the need for more nuanced privacy-preserving interventions that offer rigorous privacy guarantees without compromising factual accuracy.
Problem

Research questions and friction points this paper is trying to address.

Privacy-Hallucination Tradeoff
Differentially Private Language Models
Factual Accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Language Models
Privacy-Hallucination Tradeoff
Output Distributions