RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入RAG-Safety-Bench基准来评估检索增强型大语言模型的安全性问题,旨在明确不同因素对生成响应安全性的影响。
📝 Abstract
Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overall safety of the generated responses, when prompted for harmful or dangerous content. A clearer understanding of the mechanisms leading to this result is needed, as increasing numbers of end users turn to RAG to incorporate corporate documents and knowledge bases into LLM-based systems. We introduce RAG-Safety-Bench, a benchmark to measure the safety impact of RAG on LLM models. By removing the confounding effect of retriever quality, and cleanly separating the problem into four conditions -- non-RAG, RAG with an oracle document containing the answer to the harmful request, RAG with documents related to the harmful request but without the specific answer, and RAG with random, safe documents -- the benchmark isolates the impacts of different factors in the observed safety degradation. We report results across five open-source LLMs, showing an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not lead to downstream safety guarantees in the RAG case, and model-specific support for previous findings that even benign documents can lead to unsafe generation in retrieval-enabled systems.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Safety
Large Language Models
Harmful Content
Innovation

Methods, ideas, or system contributions that make the work stand out.

RAG-Safety-Bench
retrieval-augmented generation
safety evaluation
LLM safety
💼 Related Jobs
No related jobs found.
A
Adithiyan Rajan Indira Saravanan
Faculty of Engineering, University of Ottawa, Ottawa, Canada
Kathleen C. Fraser
Kathleen C. Fraser
University of Ottawa
computational linguisticsmachine learningnatural language processing