Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing

📅 2026-04-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the critical privacy risks posed by large language model (LLM) agents, which may inadvertently leak user-sensitive information through contextual memory during inference. Existing privacy extraction attacks often require white-box access, multiple queries, or substantial computational resources, limiting their practicality. To overcome these limitations, we propose Spore, a training-free, single-query black-box and gray-box attack framework. Its core innovation lies in an information-theoretically guided hybrid probing strategy: it constructs candidate sets under black-box settings and leverages multi-rank token analysis to enhance recovery accuracy in gray-box scenarios. Extensive experiments demonstrate that Spore significantly outperforms state-of-the-art attacks across multiple mainstream LLMs, achieving high success rates with minimal overhead and strong robustness, while effectively evading prevalent detection mechanisms and safety alignment defenses.

Technology Category

Application Category

📝 Abstract
With the wide adoption of personal AI assistants such as OpenClaw, privacy leakage in user interaction contexts with large language model (LLM) agents has become a critical issue. Existing privacy attacks against LLMs primarily target training data, while research on inference-time contextual privacy risks in LLM agent memory remains limited. Moreover, prior methods often incur high attack costs, requiring multiple queries or relying on white-box assumptions, which limits their practicality in real-world deployments. To address these issues, we propose a training-free privacy extraction attack targeting LLM agent memory, which we name \textsc{Spore}. \textsc{Spore} is compatible with both black-box and gray-box settings. In the black-box setting, \textsc{Spore} can efficiently extract a small candidate set via a single query to recover the original private information. In the gray-box setting, \textsc{Spore} allows the attacker to leverage multi-ranked tokens for more accurate and faster privacy extraction. We provide an information-theoretic analysis of \textsc{Spore} and show that it achieves high query efficiency with substantial per query information leakage. Experiments on multiple frontier LLMs show that \textsc{Spore} outperforms attack success rate over existing state-of-the-art (SOTA) schemes. It also maintains low attack cost and remains stable across different model parameter settings. We further evaluate the robustness of \textsc{Spore} against existing defense mechanisms. Our results show that \textsc{Spore} consistently bypasses both detection and strong safety alignment, demonstrating resilient performance in diverse defensive settings and real-world safety threats.
Problem

Research questions and friction points this paper is trying to address.

privacy leakage
large language models
inference-time privacy
LLM agent memory
privacy attack
Innovation

Methods, ideas, or system contributions that make the work stand out.

privacy extraction attack
training-free
inference-time probing
black-box attack
LLM agent memory
Y
Yu Cui
School of Cyberspace Science and Technology, Beijing Institute of Technology
R
Ruiqing Yue
Chengdu Institute of Computer Applications, Chinese Academy of Sciences; University of Chinese Academy of Sciences
H
Hang Fu
School of Cyberspace Science and Technology, Beijing Institute of Technology
S
Sicheng Pan
School of Cyberspace Science and Technology, Beijing Institute of Technology
Z
Zhuoyu Sun
School of Cyberspace Science and Technology, Beijing Institute of Technology
B
Baohan Huang
School of Cyberspace Science and Technology, Beijing Institute of Technology
H
Haibin Zhang
Yangtze Delta Region Institute of Tsinghua University, Zhejiang; Jiaxing Key Laboratory of Artificial Intelligence and Cyber Resilience
Cong Zuo
Cong Zuo
Beijing Institute of Technology
Cryptography
L
Licheng Wang
School of Cyberspace Science and Technology, Beijing Institute of Technology