Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

πŸ“… 2026-08-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the debate between information erasure and output suppression in large language model refusal mechanisms by introducing the concept of broken symmetry. Employing bidirectional activation patching and linear probes, we reveal a fundamental causal asymmetry: while refused answers remain linearly recoverable and can be released via single-position intervention, restoring refusal requires coordinated multi-site manipulation. These findings demonstrate that refusal does not function as a symmetric switch, thereby correcting previous overestimations of linear probes’ behavioral control capabilities. Furthermore, this work elucidates the localized nature of safety alignment, providing critical theoretical foundations and novel perspectives for auditing the safety of large language models.
πŸ“ Abstract
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we term broken symmetry. Even when a model generates a clean refusal, the correct answer remains linearly recoverable from its hidden states. Furthermore, releasing this withheld answer is a highly local operation, requiring only a single-position patch. Conversely, the reverse operation is not equally local: reimposing suppression requires broader interventions across multiple positions, and assembling a coherent refusal sequence is more difficult still. We further demonstrate that while an average answer-to-refusal displacement vector marks the geometric difference between these states, it fails to act as a reliable, reversible linear control toggle between behaviours. Taken together, our findings show that refusal does not function as a simple symmetric switch. For safety and auditing, this implies that probe recoverability can overestimate true behavioural control, and locating refusal-relevant directions does not reliably grant the ability to steer a model from answering to coherent refusal.
Problem

Research questions and friction points this paper is trying to address.

LLM Refusal
Internal Representations
Broken Symmetry
Safety Alignment
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Broken Symmetry
LLM Refusal
Activation Patching
Linear Recoverability
Intervention Locality