REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the unreliability of large reasoning models, which often suffer from reasoning errors or knowledge hallucinations. To mitigate these issues, the authors propose REIN, a novel alignment framework that uniquely integrates explicit self-reflection with an active abstention mechanism. Within a single forward pass and without requiring process supervision or external retrieval, REIN generates structured reasoning chains following a β€œthink β†’ reflect β†’ respond” paradigm and proactively abstains when confident evidence is lacking. The approach substantially reduces both types of hallucinations, achieving a 58%–72% reduction in hallucination rates compared to baseline methods across multiple benchmarks, while maintaining high coverage (86%–91%). Furthermore, among answered samples, it improves selective accuracy by 6.6–14.2%.
πŸ“ Abstract
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinct failure sources: reasoning hallucination, where flawed inference steps propagate to an incorrect conclusion, and knowledge hallucination, where the model lacks the requisite factual knowledge to answer the query. To address reasoning hallucination, we propose REIN, an alignment framework that trains LRMs to produce a structured reasoning sequence, $\texttt{<think>} $$\rightarrow$ $\texttt{<reflection>} $$\rightarrow$ $\texttt{<answer>}$, enabling explicit self-reflection before committing to a final answer. To address knowledge hallucination, REIN introduces a reward mechanism that encourages explicit abstention (e.g., "I don't know") when none of the sampled reasoning chains yields a correct answer, allowing the model to refrain from unsupported predictions. Extensive evaluations on mathematical and commonsense reasoning benchmarks show that REIN consistently improves selective accuracy, reduces incorrect-but-self-endorsed responses, and maintains high coverage compared with competitive baselines. Notably, REIN achieves these gains within a single forward pass, without requiring process supervision, inference-time controllers, external search, or multi-round critiques. Experiments on multiple backbones show that REIN reduces the hallucination proxy by $58\sim72\%$ relative to the base models while maintaining $86\sim91\%$ average coverage, and improves selective accuracy on attempted questions by $6.6\sim14.2\%$.
Problem

Research questions and friction points this paper is trying to address.

hallucination
reasoning
reliability
large reasoning models
abstention
Innovation

Methods, ideas, or system contributions that make the work stand out.

reasoning hallucination
knowledge hallucination
self-reflection
abstention alignment
selective accuracy
Z
Zhengze Huang
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
L
Luyang Yu
Fudan University
D
Di Hong
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
X
Xinzhe Huang
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Wanyu Lin
Wanyu Lin
The Hong Kong Polytechnic University
Graph LearningAI for ChemistryAI for Materials ScienceCollaborative Learning
Zhixuan Chu
Zhixuan Chu
Associate Professor, Zhejiang University; Alibaba Group; Ant Group
Zhan Qin
Zhan Qin
Researcher, Zhejiang University
Data Security and PrivacyAI Security
Tianhang Zheng
Tianhang Zheng
Zhejiang University