Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

📅 2026-05-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
The interlayer propagation mechanism of residual streams in large language models and their underlying spectral geometric structure remain poorly understood. This work treats Transformer depth as discrete time and models the residual stream as a dynamical system, integrating full Jacobian eigendecomposition, spectral geometry, graph community detection, and non-normal operator theory. It reveals, for the first time, a monotonic spectral gradient in trained models—transitioning from non-normal, rotation-dominated dynamics to near-symmetric behavior across layers. The study demonstrates that low-rank bottlenecks and the topological position of graph communities jointly govern perturbation amplification or suppression. These findings are validated across three production-scale large language models, showing that the observed spectral gradient and dimensional collapse are emergent properties of training, and that local operator types can predict perturbation propagation behavior.
📝 Abstract
Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. However, previous analyses have relied on scalar summaries or approximate linearizations, leaving the full spectral geometry of trained LLMs unknown. We perform full Jacobian eigendecomposition across three production--scale LLMs and show that training installs a monotonic spectral gradient through depth -- from non-normal, rotation-dominated early layers to near--symmetric late layers -- together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. Our experiments reveal that this gradient and the dimensional collapse are learned rather than architectural, and is largely dissolved when structured non-normality is removed. We further show that the topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type, a relationship absent at initialization. These results map a learned spectral geometry in LLMs that links perturbation propagation and compression to the network's functional topology.
Problem

Research questions and friction points this paper is trying to address.

residual stream
spectral geometry
network topology
large language models
Jacobian dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

spectral geometry
residual stream
Jacobian eigendecomposition
non-normal dynamics
network topology
🔎 Similar Papers
2023-12-17Bulletin of the American Mathematical SocietyCitations: 59
2024-07-13arXiv.orgCitations: 36