🤖 AI Summary
This work investigates how state-dependent dynamic feedback in self-attention mechanisms spontaneously gives rise to structured representations and their associated attractor geometries. By constructing a simplified normalized self-attention dynamical model and leveraging tools from statistical physics—specifically the thermodynamic limit, high-dimensional geometry, and dynamical systems theory—the study elucidates the interplay between token clustering and attention matrix formation. The authors introduce the “overlap gap” as a key order parameter governing attractor structure and identify a critical threshold in attention sharpness: only when this threshold is exceeded and intra-cluster similarity substantially surpasses inter-cluster similarity does the system undergo a dynamic condensation phase transition from an unstructured initial state, thereby self-organizing into stable clustered attractor manifolds. In this regime, cross-cluster attention decays exponentially with dimensionality, yielding a continuous spectrum of structures ranging from macroscopic clusters to microscopic fragments.
📝 Abstract
Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.