Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Reflection Steering框架,通过在激活空间中分离反射与推理,减少大型推理模型中的冗余反思,从而节省推理令牌并提高效率。
📝 Abstract
Large reasoning models often produce reasoning traces with verification, revision, and backtracking. When reflection merely re-checks established results, it wastes reasoning tokens and increases latency. Most existing reflection steering methods add a label-derived mean-difference direction across preset layers, but its entanglement with reasoning and length signals destabilizes the accuracy-efficiency trade-off. In this paper, we propose Reflection Steering, a training-free framework for controlling reflection-associated computation within LLMs by disentangling reflection-related activations from general reasoning. Specifically, we contrast reflective and non-reflective hidden states at each LLM layer, denoise the resulting reflection directions with PCA, and orthogonalize them against general-reasoning directions. To limit downstream amplification from early-layer interventions, we calibrate each layer across multiple intervention strengths on a small set, retain only stable layers, and apply bounded projection removal to their residual-stream activations. We conduct extensive experiments across two public benchmarks and three open-weight LLMs against state-of-the-art activation-steering baselines. Results show that Reflection Steering reduces reasoning tokens by 16.9% on average across six matched settings. Besides, our method further introduces a bounded reflection intervention-strength parameter $α$, enabling deployment-time adjustment to balance token savings, accuracy, and generation stability.
Problem

Research questions and friction points this paper is trying to address.

reflection
reasoning
token-efficient inference
large reasoning models
activation space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reflection Steering
disentangling reflection
activation space
token-efficient inference
bounded projection removal
🔎 Similar Papers
Jiarui Hu
Jiarui Hu
Zhejiang University
Computer Vision Robotics Computer Graphics
Zhiyuan Wen
Zhiyuan Wen
The Hong Kong Polytechnic University
NLP
X
Xiaoyun Liu
Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
J
Jiaxing Shen
School of Data Science, Lingnan University, Hong Kong SAR, China
Y
Yu Yang
Centre for Learning, Teaching and Technology, The Education University of Hong Kong, Hong Kong SAR, China