VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy

📅 2025-10-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the security vulnerability of black-box deployed conversational AI systems. We propose VortexPIA, an indirect prompt injection attack that requires no access to or modification of the system’s internal prompts, yet effectively induces large language models (LLMs) to proactively solicit users’ multi-category private information. Methodologically, VortexPIA leverages low-overhead false-memory embedding and adversarial input design to achieve targeted privacy extraction. It is the first approach to enable token-efficient, batched induction of diverse sensitive data categories in fully black-box settings. Evaluated across six mainstream LLMs and four benchmark datasets, VortexPIA achieves state-of-the-art performance—reducing token consumption by up to 42% and improving request success rates by 3.1× over prior methods. Furthermore, it demonstrates high stealth, practical efficacy, and robust evasion of existing defense mechanisms in multiple open-source real-world applications.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) have been widely deployed in Conversational AIs (CAIs), while exposing privacy and security threats. Recent research shows that LLM-based CAIs can be manipulated to extract private information from human users, posing serious security threats. However, the methods proposed in that study rely on a white-box setting that adversaries can directly modify the system prompt. This condition is unlikely to hold in real-world deployments. The limitation raises a critical question: can unprivileged attackers still induce such privacy risks in practical LLM-integrated applications? To address this question, we propose extsc{VortexPIA}, a novel indirect prompt injection attack that induces privacy extraction in LLM-integrated applications under black-box settings. By injecting token-efficient data containing false memories, extsc{VortexPIA} misleads LLMs to actively request private information in batches. Unlike prior methods, extsc{VortexPIA} allows attackers to flexibly define multiple categories of sensitive data. We evaluate extsc{VortexPIA} on six LLMs, covering both traditional and reasoning LLMs, across four benchmark datasets. The results show that extsc{VortexPIA} significantly outperforms baselines and achieves state-of-the-art (SOTA) performance. It also demonstrates efficient privacy requests, reduced token consumption, and enhanced robustness against defense mechanisms. We further validate extsc{VortexPIA} on multiple realistic open-source LLM-integrated applications, demonstrating its practical effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Attacks extract private data from LLMs in black-box settings
Indirect prompt injection misleads models to request sensitive information
Evaluates attack efficiency across multiple LLMs and real applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect prompt injection attack in black-box settings
Uses token-efficient false memory data injection
Flexibly targets multiple categories of sensitive information
Y
Yu Cui
Beijing Institute of Technology
S
Sicheng Pan
Beijing Institute of Technology
Y
Yifei Liu
Beijing Institute of Technology
H
Haibin Zhang
Yangtze Delta Region Institute of Tsinghua University, Zhejiang
Cong Zuo
Cong Zuo
Beijing Institute of Technology
Cryptography