Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了激活导向方法在控制语言模型行为时的效果,发现该方法主要使模型趋向于其固有偏好行为而非预期行为,并通过24种行为和10个模型验证了这一现象。
📝 Abstract
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.
Problem

Research questions and friction points this paper is trying to address.

Activation Steering
Behavior Interference
Language Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

activation steering
behavior directions
model defaults
modular control
🔎 Similar Papers
S
Srikanth Malla
Samsung Semiconductor US
Chiho Choi
Chiho Choi
Sr. Staff ML Engineer, Samsung DSA
Machine LearningComputer Vision
J
Joon Hee Choi
Samsung Semiconductor US