🤖 AI Summary
The interlayer propagation mechanism of residual streams in large language models and their underlying spectral geometric structure remain poorly understood. This work treats Transformer depth as discrete time and models the residual stream as a dynamical system, integrating full Jacobian eigendecomposition, spectral geometry, graph community detection, and non-normal operator theory. It reveals, for the first time, a monotonic spectral gradient in trained models—transitioning from non-normal, rotation-dominated dynamics to near-symmetric behavior across layers. The study demonstrates that low-rank bottlenecks and the topological position of graph communities jointly govern perturbation amplification or suppression. These findings are validated across three production-scale large language models, showing that the observed spectral gradient and dimensional collapse are emergent properties of training, and that local operator types can predict perturbation propagation behavior.
📝 Abstract
Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. However, previous analyses have relied on scalar summaries or approximate linearizations, leaving the full spectral geometry of trained LLMs unknown. We perform full Jacobian eigendecomposition across three production--scale LLMs and show that training installs a monotonic spectral gradient through depth -- from non-normal, rotation-dominated early layers to near--symmetric late layers -- together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. Our experiments reveal that this gradient and the dimensional collapse are learned rather than architectural, and is largely dissolved when structured non-normality is removed. We further show that the topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type, a relationship absent at initialization. These results map a learned spectral geometry in LLMs that links perturbation propagation and compression to the network's functional topology.