Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型推理模型中中间推理过程和最终答案的安全性问题,提出了一种基于段落感知列表对齐的方法SaLT-DPO,通过独立评分与联合安全一致性正则化来提高安全性。
📝 Abstract
Large Reasoning Models (LRMs) pose a dual-surface safety challenge: both intermediate reasoning traces and final answers can contain harmful content. Existing alignment methods often operate at the whole-response level, allowing unsafe reasoning to be masked by a safe-looking final answer. We propose Segment-aware Listwise Target DPO (SaLT-DPO), which addresses this gap through three mechanisms: (1) segment-aware listwise alignment that decomposes responses into reasoning and answer segments, independently scores each segment's safety, and aligns length-normalized segment rewards with soft target distributions over multiple candidates; (2) joint safety coherence regularization that applies a weakest-link principle to promote safety consistency across both segments; and (3) utility anchoring on benign prompts to mitigate over-refusal and reasoning degradation. Experiments on three LRMs show that SaLT-DPO consistently reduces unsafe rates for both reasoning and answer segments while mitigating degradation in benign compliance and preserving general reasoning performance. Ablation studies demonstrate the complementary contributions of its components.
Problem

Research questions and friction points this paper is trying to address.

Large Reasoning Models
safety challenge
alignment methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Segment-aware Listwise Alignment
Safety Coherence Regularization
Utility Anchoring
💼 Related Jobs
No related jobs found.
JungMin Yun
JungMin Yun
Chung-Ang University
Deep LearningNatural Language ProcessingHuman-Centered AI
J
Junehyoung Kwon
Department of Artificial Intelligence, Chung-Ang University
H
Hayeong Ryu
Department of Artificial Intelligence, Chung-Ang University
B
Byeonggeuk Lim
Department of Artificial Intelligence, Chung-Ang University
H
Hoejoon Kwon
Department of Artificial Intelligence, Chung-Ang University
YoungBin Kim
YoungBin Kim
Chung-Ang University
Machine Learning