Guaranteeing Faithful Evidence Extraction in Speculative Retrieval-Augmented Generation

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型在信息检索中产生幻觉和忠实性错误的问题,提出了一种约束混合解码方法,确保生成的答案与检索到的证据完全一致。
📝 Abstract
Large Language Models (LLMs) are increasingly used as interfaces for information retrieval, but they remain prone to hallucinations and faithfulness errors, in which the generated answers diverge from the retrieved evidence. While Retrieval-Augmented Generation (RAG) and recent hybrid or semi-extractive approaches mitigate this issue, they do not guarantee that quoted or extracted spans are verbatim from the retrieved context. This limitation can have severe consequences in safety-critical domains, where answers must exactly match certified documentation. We introduce Constrained Hybrid Decoding (CHyD), a novel faithfulness-first paradigm for speculative RAG. While traditional speculative decoding is optimized for inference speed, CHyD repurposes this architecture to ensure faithful verbatim evidence extraction when the extraction mode is correctly triggered. Our approach enforces hard decoding constraints that restrict generation to continuous spans present in the retrieved documents. This design provides a robust but straightforward guarantee: any explicitly quoted span in the output appears verbatim in the provided context. We evaluate our method across state-of-the-art LLMs on diverse abstractive, extractive, and semi-extractive QA benchmarks, including technical datasets motivated by aircraft maintenance. Results show that existing hybrid methods frequently hallucinate quoted spans, with exact extraction accuracy dropping below 40% in technical domains. In contrast, our approach achieves near-perfect extraction faithfulness regardless of the model used. Although enforcing hard constraints introduces a trade-off with fluency-oriented metrics, our method improves exact answer correctness and remains competitive overall, highlighting its suitability for safety-critical information retrieval applications.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
faithfulness errors
Retrieval-Augmented Generation
verbatim extraction
safety-critical domains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Hybrid Decoding
faithfulness-first
speculative RAG
verbatim evidence extraction
safety-critical domains
💼 Related Jobs
No related jobs found.
Q
Quentin Signé
Université de Toulouse - IRIT UMR 5505; Airbus Protect
Mohand Boughanem
Mohand Boughanem
Université de Toulouse - IRIT UMR 5505
J
Jose Moreno
Université de Toulouse - IRIT UMR 5505
T
Thiziri Belkacem
Airbus Protect