Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks

📅 2025-09-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Factuality errors in large language model (LLM) question answering—primarily caused by hallucination—are inadequately captured by existing entropy-based semantic uncertainty estimation methods, which suffer from sampling noise and clustering instability induced by variable-length outputs. Method: We propose Semantic Reconstruction Entropy (SRE), a robust uncertainty quantification framework that first augments input-side semantic diversity via faithful paraphrasing, then applies energy-driven progressive hybrid clustering in semantic space to enable adaptive, length-agnostic grouping and uncertainty estimation. Contribution/Results: SRE eliminates sensitivity to output length and sampling strategy. Extensive experiments on SQuAD and TriviaQA demonstrate that SRE significantly outperforms state-of-the-art baselines in hallucination detection accuracy, stability under distributional shift, and cross-dataset generalization—achieving consistent improvements across all metrics.

Technology Category

Application Category

📝 Abstract
Reliable question answering with large language models (LLMs) is challenged by hallucinations, fluent but factually incorrect outputs arising from epistemic uncertainty. Existing entropy-based semantic-level uncertainty estimation methods are limited by sampling noise and unstable clustering of variable-length answers. We propose Semantic Reformulation Entropy (SRE), which improves uncertainty estimation in two ways. First, input-side semantic reformulations produce faithful paraphrases, expand the estimation space, and reduce biases from superficial decoder tendencies. Second, progressive, energy-based hybrid clustering stabilizes semantic grouping. Experiments on SQuAD and TriviaQA show that SRE outperforms strong baselines, providing more robust and generalizable hallucination detection. These results demonstrate that combining input diversification with multi-signal clustering substantially enhances semantic-level uncertainty estimation.
Problem

Research questions and friction points this paper is trying to address.

Detecting fluent but factually incorrect hallucinations in LLM question answering
Overcoming limitations of noisy sampling in semantic uncertainty estimation methods
Addressing unstable clustering issues with variable-length answer formulations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Input-side semantic reformulations create faithful paraphrases
Progressive energy-based hybrid clustering stabilizes semantic grouping
Combining input diversification with multi-signal clustering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chaodong Tong
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Q
Qi Zhang
China Industrial Control Systems Cyber Emergency Response Team, Beijing, China
L
Lei Jiang
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Y
Yanbing Liu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
N
Nannan Sun
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
W
Wei Li
China Industrial Control Systems Cyber Emergency Response Team, Beijing, China