Representational alignment yields generalizable safety in language models

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文探讨了通过优化语言模型中的潜在表示与人类道德判断的一致性,来提高模型在面对对抗性或不熟悉形式的有害意图时的安全性和鲁棒性。
📝 Abstract
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.
Problem

Research questions and friction points this paper is trying to address.

large language models
moral categorization
adversarial conditions
prototype theory
alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

representational similarity optimization
moral categorization
adversarial robustness
🔎 Similar Papers
2024-06-20arXiv.orgCitations: 26
Lingyu Li
Lingyu Li
Shanghai Jiao Tong University
Active inferenceArtificial Intelligencephilosophy
Y
Yan Teng
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Y
Yingchun Wang
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Xia Hu
Xia Hu
Google DeepMind
Deep LearningMachine LearningMultimodal