🤖 AI Summary
This study addresses the unclear mechanisms by which redundancy and diversity in retrieved documents affect Retrieval-Augmented Generation (RAG) quality. Through controlled experiments on the FictionalQA synthetic dataset, we effectively eliminate interference from parametric knowledge. Our findings demonstrate that genre diversity within the retrieval set is a critical factor for improving generation accuracy, whereas information repetition or paraphrasing yields no significant gains. Specifically, introducing diversity increases answer correctness by 17% to 47%. By elucidating how structural characteristics of the retrieved context influence RAG performance, this work provides robust empirical evidence and theoretical guidance for optimizing retrieval strategies.
📝 Abstract
In RAG, while the retriever typically ranks documents by their individual relevance to the query, the generator instead produces an answer based on the retrieved documents as a whole. This paper investigates how redundancy and diversity from the retrieved document set impact the generator in terms of answer correctness.
Previous work has provided a mix of findings: some showing that redundancy improves generation by reinforcing relevant information, others that LLM-based paraphrasing of the same content may be beneficial. Many of these studies did not control for confounding factors like whether the documents contained the exact answer or not, and if parametric knowledge plays a role. We conduct a carefully controlled experiment investigating three key scenarios of retrieved document sets: 1) Duplicate (exact copies of the same document), 2) Paraphrased (LLM rephrased versions of one document) and 3) Diverse (documents from different genres each containing relevant information in different forms). We control for which documents contain the answer in exact match or rephrased form.
Evaluation is done with FictionalQA, a synthetic, fictional question-answer dataset that ensures the LLM generator prior knowledge cannot answer the question; the answer must come from retrieved documents. We show that duplicate redundancy and LLM paraphrasing does not significantly improve answer correctness. However, providing diverse documents is highly beneficial, improving answer correctness by 17%-47%. We further show this improvement is driven by diverse forms of document genre (news, blogs, etc.) alone and not a consequence of more relevant answer being available to generator. Our findings help to direct more attention to how new retrieval methods might improve RAG by catering to the generator preference for diversity in retrieval results.