Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过分解安全调整数据集中的响应为拒绝声明和解释理由两部分,以减少语言模型中因依赖表面线索而产生的错误拒绝问题。
📝 Abstract
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues. In contrast, training solely on rationales reduces false refusals while maintaining a comparable level of safety performance. Rationale-Only benefits also appear in our ICL configuration and remain compatible with the evaluated inference-time mitigation methods. The results emphasize the necessity of precisely curated, fine-grained safety supervision datasets and outline directions for constructing aligned agents that better reconcile helpfulness with safety.
Problem

Research questions and friction points this paper is trying to address.

safety
helpfulness
false refusals
language models
harmful queries
Innovation

Methods, ideas, or system contributions that make the work stand out.

safety-tuning
rationale
false refusals
language models
💼 Related Jobs
No related jobs found.
M
Minji Kim
Graduate School of Artificial Intelligence, POSTECH
Hyounghun Kim
Hyounghun Kim
POSTECH
NLPMultimodal Learning