🤖 AI Summary
This study investigates the cross-layer dynamical mechanisms of the residual stream (RS) in Transformer models. We model the RS as a continuous dynamical system—introducing dynamical systems theory from neuroscience into large-model interpretability research for the first time. Using orbital stability analysis, dimensionality-reduction visualizations (PCA/t-SNE), and large-scale activation statistics, we identify three key properties: (i) strong cross-layer continuity in the RS; (ii) inter-layer acceleration of evolution, exponential growth in activation density, and unstable periodic orbits; and (iii) curved, attractor-like trajectories in low-dimensional embedding space. Collectively, these findings reveal an underlying dynamical structure in the RS characterized by coexisting stable and unstable regimes. Our work establishes the first dynamical-systems-based theoretical framework for large-model interpretability, grounded in empirical evidence—thereby laying foundational groundwork for an AI neuroscience paradigm.
📝 Abstract
As artificial intelligence models have exploded in scale and capability, understanding of their internal mechanisms remains a critical challenge. Inspired by the success of dynamical systems approaches in neuroscience, here we propose a novel framework for studying computations in deep learning systems. We focus on the residual stream (RS) in transformer models, conceptualizing it as a dynamical system evolving across layers. We find that activations of individual RS units exhibit strong continuity across layers, despite the RS being a non-privileged basis. Activations in the RS accelerate and grow denser over layers, while individual units trace unstable periodic orbits. In reduced-dimensional spaces, the RS follows a curved trajectory with attractor-like dynamics in the lower layers. These insights bridge dynamical systems theory and mechanistic interpretability, establishing a foundation for a"neuroscience of AI"that combines theoretical rigor with large-scale data analysis to advance our understanding of modern neural networks.