Can Activation Steering Capture Multidimensional Authorship Style?

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过结构化对比提示在激活空间中构建丰富风格表示,提出无训练框架A3S,解决多维度写作风格控制问题。
📝 Abstract
Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it remains unclear whether it can handle the multidimensional and hard-to-define nature of authorship style. We ask whether structured contrastive prompting along rhetorically-motivated dimensions can construct rich style representations directly in activation space, bypassing the need for natural language style descriptors or dedicated training. We find that the resulting directions share a common authorship backbone while conflicting on aspect-specific residuals that carry genuine stylistic signal, explaining why naive aggregation fails. We operationalize this in Aspect-Aware Activation Steering (A3S), a training-free framework that merges per-aspect contrastive directions with interference-aware aggregation and tunes steering strength per instance. A3S improves authorship style transfer where it is genuinely multi-aspect, outperforms a trained baseline in preference evaluations on out-of-domain benchmarks, and keeps target-exemplar overlap consistently low.
Problem

Research questions and friction points this paper is trying to address.

activation steering
authorship style
multidimensional
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Steering
Contrastive Prompting
Authorship Style Transfer
A3S
Interference-aware Aggregation
🔎 Similar Papers