Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对检索增强生成模型中的安全性和鲁棒性问题,通过定义威胁模型、分类攻击目标,并从管道视角审视防御措施来提升模型的可靠性。
📝 Abstract
Retrieval-Augmented Generation (RAG) enhances large language models by grounding outputs in external knowledge, improving factuality and reducing hallucinations. At the same time, the retrieval-augmented pipeline introduces new robustness and security risks, including corpus poisoning, backdoor attacks, privacy leakage, and fairness violations. Despite rapid progress in this area, existing surveys remain limited in their treatment of attacker objectives, threat models, and stage-specific defenses across the full RAG pipeline. This survey presents a unified and pipeline-aware overview of RAG robustness. We formalize threat models over the corpus, retriever, and generator, and organize attacks into three main objectives: accuracy, privacy, and fairness. We further review defenses from a pipeline-aware perspective, covering the retrieval, rerank, generation, and traceback stages. In addition, we summarize robustness benchmarks and explainability methods for more deeply evaluating and explaining RAG robustness.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
robustness
security risks
threat models
defenses
Innovation

Methods, ideas, or system contributions that make the work stand out.

pipeline-aware
threat models
accuracy, privacy, and fairness
defenses
robustness benchmarks