Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection
This work addresses the vulnerability of existing retrieval-augmented intrusion detection systems (RAG-IDS) to knowledge poisoning and prompt injection attacks, which degrade performance and induce high false-positive rates. To mitigate these threats, the authors propose a novel three-tier multi-agent RAG-IDS framework that integrates soft trust scoring, label embedding consistency checking (LECC), and prompt sanitization mechanisms to establish a robust retrieval-boundary defense. Experimental results demonstrate substantial improvements in system robustness: under 30% knowledge poisoning on the CIC-UNSW-NB15 dataset, the approach achieves a classification performance recovery rate of 0.57; under prompt injection attacks, the label-flipping success rate is reduced to 0.6–2.4%, significantly outperforming single-document retrieval baselines, which exhibit rates of 35–55%.