CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CLEAR方法,通过轻量级隐藏状态门控制安全低秩适配器的激活强度,以减少有害输出同时保持良性输入性能,改善了LLM的安全性与实用性之间的权衡。
📝 Abstract
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E}nt \textbf{A}dapter \textbf{R}outing (CLEAR), a conditional safety adaptation framework that uses a lightweight hidden-state gate to continuously control the activation strength of a safety low-rank adapter. CLEAR aims to reduce harmful completions while avoiding unnecessary changes to the frozen backbone that could degrade performance on benign prompts. Experiments on widely used safety and utility benchmarks show that CLEAR improves robustness on HarmBench while reducing the utility degradation observed with globally applied safety tuning such as SFT or standard low-rank adaptation (LoRA). On Llama-3-8B-Instruct, CLEAR reduces HarmBench ASR from 32.3\% to 0.5\%, while retaining most of the base model's utility and achieving up to 7.1 percentage points higher GSM8K accuracy than globally applied SFT or LoRA. These results suggest that CLEAR is a promising mechanism for improving the safety--utility trade-off in LLM alignment.
Problem

Research questions and friction points this paper is trying to address.

large language models
safety
utility
adaptation
trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Safety Adaptation
Latent Adapter Routing
Safety Low-Rank Adapter
Hidden-State Gate
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chengxiao Wang
Siebel School of Computing and Data Science, University of Illinois at Urbana-Champaign
E
Enyi Jiang
Siebel School of Computing and Data Science, University of Illinois at Urbana-Champaign; Computer Science, Stanford University
Xiaojing Liao
Xiaojing Liao
University of Illinois Urbana-Champaign
Cybersecurity
Sanmi Koyejo
Sanmi Koyejo
Assistant Professor, Stanford University
Machine LearningHealthcare AINeuroinformatics