Inferring Hidden User Models from the Behavior of Personalized LLM Agents

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种名为UMPeek的黑盒攻击方法,通过假设引导的自适应探测来推断隐藏的用户模型,从而揭示了即使在源记录和后端状态不可访问的情况下,个性化LLM代理仍可能泄露隐私信息的问题。
📝 Abstract
Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more privacy-preserving because direct memory-extraction attacks lose the text they target. Yet we argue that user models expose a new attack surface because an attacker can still recover the private information from the personalized choices they shape, even when source records and backend state remain inaccessible. We therefore introduce UMPeek, a black-box attack based on hypothesis-guided adaptive probing to infer such hidden user model. It forms hypotheses from choices left open by a request, switches among ordinary follow-up tasks, and retains only claims supported and not contradicted by visible behavior. We conduct an extensive benchmark evaluation across diverse personalization tasks and user-model backends against existing attacks. We further validate UMPeek in real-world systems using information confirmed to be retained, and we evaluate defenses against its adaptive probing. Overall, UMPeek outperforms existing attacks in both benchmark and real-world comparisons and continues to recover user information under response-level defenses, showing that keeping records and backend state inaccessible does not guarantee semantic privacy when retained information shapes visible behavior.
Problem

Research questions and friction points this paper is trying to address.

user models
privacy-preserving
attack surface
personalized choices
hidden information
Innovation

Methods, ideas, or system contributions that make the work stand out.

User Models
Hypothesis-Guided Adaptive Probing
Semantic Privacy
🔎 Similar Papers
No similar papers found.
H
Haoyang Li
The Hong Kong Polytechnic University, Hong Kong SAR
Y
Yaxin Xiao
The Hong Kong Polytechnic University, Hong Kong SAR
Qingqing Ye
Qingqing Ye
Assistant Professor, The Hong Kong Polytechnic University
data privacy and securityadversarial machine learning
Huadi Zheng
Huadi Zheng
Unknown affiliation
Voice TechnologyInformation Security
H
Haibo Hu
The Hong Kong Polytechnic University, Hong Kong SAR