Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of reliably distinguishing correct from incorrect reasoning in large language models without relying on superficial shortcuts. It introduces the first approach that models reasoning errors as region- and direction-specific signals within residual streams and proposes a three-stream detector that integrates residual trajectory dynamics, vector-quantized coarse-grained regional information, and fine-grained directional cues from normalized multi-layer states to reconstruct rich contextual representations. By moving beyond methods limited to token-level shifts or single-layer probing, the proposed framework achieves up to a 12% improvement in selection accuracy over existing shift-based methods and a 21% gain over single-layer baselines on unseen reasoning benchmarks, while consistently outperforming competing probes in factual completion and verification tasks.
📝 Abstract
As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek this signal in layerwise residual-stream displacements, which capture how representations change while attenuating some stable, token-specific information. However, displacement omits the state from which an update originates, whereas restoring the full state risks reintroducing shortcut-prone information. We identify this trade-off and propose a three-stream detector that combines motion with two restricted views of location. A coarse region reader based on vector quantization and a fine direction reader over normalized multi-layer states. This design restores enough state context to interpret the motion without returning to full-state probing. On reasoning benchmarks unseen during training, our method improves selection accuracy by up to 12% over the displacement-only state of the art and 21% over single-layer probing baselines. Although trained only on reasoning benchmarks, it also reads factual completion and fact verification, ahead of every detector we compare against, which places the signal on correctness rather than on a kind of reasoning. Ablations further show that motion, region, and direction provide complementary signals. These results suggest that reasoning validity is better read from state-conditioned motion than from either static states or decontextualized trajectories alone.
Problem

Research questions and friction points this paper is trying to address.

reasoning errors
residual-stream trajectory
large language models
reasoning validity
representation probing
Innovation

Methods, ideas, or system contributions that make the work stand out.

residual-stream trajectory
reasoning error detection
three-stream detector
state-conditioned motion
vector quantization
🔎 Similar Papers
2024-10-03International Conference on Learning RepresentationsCitations: 28
💼 Related Jobs
No related jobs found.
H
Hamed Damirchi
Australian Institute for Machine Learning, Adelaide University, Naval Group Pacific
I
Ignacio Meza De la Jara
Australian Institute for Machine Learning, Adelaide University, Naval Group Pacific
D
Damith Ranasinghe
Adelaide University, Naval Group Pacific
Yuhang Liu
Yuhang Liu
The University of Adelaide
Representation LearningLLMsLatent Variable ModelsResponsible AI
J
Javen Shi
Australian Institute for Machine Learning, Adelaide University, Responsible AI Research Centre, Australia