ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning Attacks

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ToxicRAG,通过单文档注入误导性信息来攻击检索增强生成系统,影响其输出准确性。
📝 Abstract
Retrieval-Augmented Generation (RAG) can ground large language model (LLM) outputs in external evidence, but it also exposes the system to knowledge poisoning. Representative attacks use multiple injected documents or templates that directly assert a target answer. We present ToxicRAG, a one-document-per-target attack that expresses misinformation as a coherent knowledge-update narrative. The generated document first acknowledges the previously accepted answer, introduces fabricated events that appear to invalidate it, and then attributes the attacker-selected answer to a set of purported authorities. An answer-focused self-validation loop optionally revises a candidate when a surrogate language model does not reproduce the target answer. We evaluate the attack on 100 target questions from each of Natural Questions, HotpotQA, and MS-MARCO, using four victim LLMs and four dense retrievers. In the sampled-corpus setting reported in this paper, ToxicRAG obtains ASRs between 0.61 and 0.91 across the twelve dataset--model combinations. It matches or exceeds the strongest evaluated baseline in every combination, with margins ranging from 0 to 11 percentage points. These results show that narrative-form poisoned documents can remain influential under the evaluated RAG configurations and motivate further study of factual consistency and source provenance in RAG systems.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
knowledge poisoning
narrative-form poisoned documents
factual consistency
source provenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single-Shot Knowledge Poisoning
Coherent Knowledge-Update Narrative
Answer-Focused Self-Validation Loop
💼 Related Jobs
No related jobs found.
H
Haozhe Lu
School of Software and Microelectronics, Peking University, Beijing, China
Jiaqi Li
Jiaqi Li
Huazhong University of Science and Technology
Computer VisionDepth Estimation
X
Xinyuan Zhu
College of Cryptology and Cyber Science, Nankai University, Tianjin, China
Xiang Li
Xiang Li
Nankai University
Network SecurityIPv6 SecurityDNS Security