Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示深度推理可能导致对齐崩溃问题,并提出使用对齐损失率(ALR)量化此现象,通过注意力稀释解释原因,并提出推理残差对齐(RRA)方法缓解。
📝 Abstract
The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended reasoning remains underexplored. This paper challenges the prevailing view by revealing a critical vulnerability: Deep Reasoning May Induce Alignment Collapse. To rigorously quantify this phenomenon, we propose the Alignment Loss Rate (ALR) metric. Our experiments demonstrate that as reasoning depth increases, ALR rises significantly, indicating a severe degradation in model robustness against external perturbations. Capitalizing on this instability, a novel jailbreaking paradigm, Reasoning Trap (RT), is proposed. RT induces the model into extended reasoning to amplify the impact of adversarial attacks, leading to a sharp decline in safety capabilities. To elucidate the mechanism behind this collapse, we identify Attention Dilution as the root cause, arising from the competition for attention between the extended reasoning process and the original input. To mitigate this, Reasoning Residual Alignment (RRA), a lightweight defense strategy that dynamically re-emphasizes the input via residual connections integrated with the reasoning process.
Problem

Research questions and friction points this paper is trying to address.

Deep Reasoning
Alignment Collapse
Large Reasoning Models
Chain-of-Thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

Alignment Loss Rate (ALR)
Reasoning Trap (RT)
Attention Dilution
Reasoning Residual Alignment (RRA)
Y
Yu-Hang Wu
Shanghai University of Engineering Science
Y
Yu-Jie Xiong
Shanghai University of Engineering Science
H
Henghua Zhang
Shanghai University of Engineering Science
B
Bairui Zhang
The Hong Kong University of Science and Technology
Jia-Chen Zhang
Jia-Chen Zhang
Shanghai University of Engineering Science
large language models
S
Shaohua Li
A* STAR