Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering

๐Ÿ“… 2026-08-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡้€š่ฟ‡้ปŽๆ›ผไผ˜ๅŒ–ๅญฆไน ๅ‚ๆ•ฐ้ซ˜ๆ•ˆ็š„ๆ—‹่ฝฌๅ˜ๆขๆฅๆŽงๅˆถๅคง่ฏญ่จ€ๆจกๅž‹็š„ๆ‹’็ป่กŒไธบ๏ผŒๆ— ้œ€่พ…ๅŠฉ็ป“ๆž„๏ผŒๅฎž้ชŒ่ฏๆ˜Žไบ†่ฏฅๆ–นๆณ•็š„ๆœ‰ๆ•ˆๆ€งใ€‚
๐Ÿ“ Abstract
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.
Problem

Research questions and friction points this paper is trying to address.

Activation Steering
Refusal Behavior
Riemannian Optimization
Large Language Models (LLMs)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Riemannian optimization
Activation steering
Parameter-efficient rotations
LLM control
๐Ÿ”Ž Similar Papers