Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Suan算法,通过在梯度级别直接优化偏好,解决大型语言模型中安全对齐问题,同时保持响应的实用性。
📝 Abstract
Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Safety Alignment
Post-trained Variants
Over-refusal
General Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference Optimization
Gradient Level Formulation
Safety Alignment
Robust Training Dynamics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Oleksandr Cherednichenko
Department of Mathematics and Mathematical Statistics, Integrated Science Lab, Umeå University, Umeå, Sweden
R
Roman Klypa
Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK, 38000 Grenoble, France