Signatures of Steerability in Activation Space of Language Models

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比表示来控制语言模型行为的有效性问题,使用简单的分离度量方法评估激活空间的可操纵性。
📝 Abstract
Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steerability of language models across diverse settings, even after controlling for layers and dataset effects. Beyond prediction, we provide evidence from a synthetic superposition experiment that separation metrics are strongly correlated with alignment between the empirical and true feature direction. Our results suggest that simple separability statistics can serve as practical diagnostics for when steering vectors are likely to work.
Problem

Research questions and friction points this paper is trying to address.

steering
language models
activation space
contrastive representations
dataset dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

separation metrics
steerability
language models
P
Prajjwal Bhattarai
New York University Abu Dhabi, Abu Dhabi, UAE; New York University, Tandon School of Engineering, Brooklyn, NY, USA
Tuka Alhanai
Tuka Alhanai
New York University Abu Dhabi
machine learningcomputer sciencesignal processing