Compositional Multilingual and Behavioral Attribute Steering

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在大型语言模型中,通过无训练的属性向量加性组合方法,对语言、越狱和简洁性进行控制的有效性,并分析了其几何特性。
📝 Abstract
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attribute steering vectors can preserve the intended steering effect of each attribute, across four instruction-tuned models from two model families and two size scales. We find that single-attribute steering is reliable for all three attributes, but only within an appropriate combination of intervention layer and steering strength, with abstract behaviors (jailbreak, conciseness) favoring middle layers and language favoring earlier layers. We show that additive composition of two attribute vectors succeeds in steering both attributes simultaneously when each is injected at its own best-performing layer, and that this partially extends to three simultaneously composed attributes, addressing an inconsistency left open by prior work on training-free composition. We further analyze the geometric properties of these steering vectors, finding that they are approximately orthogonal in the residual stream, consistent with their compositional behavior.
Problem

Research questions and friction points this paper is trying to address.

compositional steering
large language models
attribute vectors
instruction-tuned models
jailbreak
Innovation

Methods, ideas, or system contributions that make the work stand out.

additive composition
attribute steering vectors
intervention layer
orthogonality
residual stream
🔎 Similar Papers
No similar papers found.