When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究良性微调下大语言模型安全对齐的脆弱性,提出基于Fisher几何的解释,并探讨了LoRA和ASRAM在缓解早期崩溃中的作用。
📝 Abstract
Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher-geometric explanation: safety Fisher is low-rank, and alignment makes the safety geometry flatter while preserving an output-routing pathway. After 100 benign fine-tuning examples, this pathway is selectively re-sharpened in output-side MLP modules, explaining the asymmetric fragility: safety can collapse to high attack success rates, while general utility degrades mildly. The routing view also explains why few safety examples can restore refusal behavior, indicating that internal safety-relevant representations are preserved. Finally, we show that LoRA and ASAM mitigate early collapse by suppressing output-side sharpness, but their protection weakens at larger fine-tuning scales. Overall, safety failure is best understood as a disruption of a low-rank output-routing mechanism
Problem

Research questions and friction points this paper is trying to address.

benign fine-tuning
safety alignment
refusal behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisher-geometric explanation
low-rank safety Fisher
output-routing pathway
asymmetric fragility
LoRA and ASAM
🔎 Similar Papers
No similar papers found.