Vulnerabilities in Personalization: Assessing Health Privacy Risks in ChatGPT Logs and Memory

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析179,057次对话,评估了ChatGPT中个人健康信息泄露的风险,并发现系统在未获得明确同意的情况下隐式提取和合成用户健康数据的问题。
📝 Abstract
As conversational LLMs become deeply embedded in daily life, users frequently disclose sensitive personal health information during routine interactions. We present a large-scale computational audit analyzing 179,057 conversations across India, Nigeria, Brazil, and Pakistan (N = 1,057) to evaluate personal health disclosures and background memory synthesis in ChatGPT. We find that 21.31% of audited conversations contain personal health data, with 3.62% posing high-to-extreme privacy risks involving stigmatized conditions, direct identifiers, and precise locations. When evaluating the memory entries of ChatGPT, we uncover a stark disconnect between corporate framing and system behavior: over 95% of profile entries are implicitly extracted without explicit user prompts or consent. Furthermore, background memory synthesis selectively condenses temporary, symptom-level disclosures into permanent diagnostic traits, stripping contextual integrity and amplifying re-identification risks. We conclude with sociotechnical design guidelines to restore user agency and consent-driven boundaries in stateful AI systems.
Problem

Research questions and friction points this paper is trying to address.

Health Privacy
Personalization
ChatGPT
Memory Synthesis
Re-identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

personal health data
privacy risks
background memory synthesis
user consent
sociotechnical design
🔎 Similar Papers