How Much Do Legal RAG Systems Still Hallucinate?

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical risks posed by hallucinations in legal Retrieval-Augmented Generation (RAG) systems through the first multi-system, fine-grained empirical analysis on GDPR and French Civil Law corpora. By constructing expert-validated benchmarks and false-premise test sets, combined with claim-level and answer-level evaluations, this work reveals distinct hallucination patterns across varying query types and user profiles. Results indicate that hallucinations remain pervasive, with significant performance disparities ranging from 10% to 50% among systems. Notably, false-premise queries substantially exacerbate hallucination rates. These findings provide essential empirical evidence for enhancing the reliability and trustworthiness of legal AI applications in high-stakes domains.
📝 Abstract
Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.
Problem

Research questions and friction points this paper is trying to address.

Legal RAG
Hallucination
False-premise questions
Ungrounded answers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Legal RAG
Hallucination Evaluation
Fine-grained Analysis
False-premise Questions
Claim-level Assessment