Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work uncovers a unified mechanism underlying both staircase-like learning dynamics—characterized by plateaus followed by abrupt loss drops—and smooth power-law loss decay in neural network training. Leveraging the permutation symmetry of units within network layers, the authors propose a unified model centered on the quadratic form $\mathrm{Tr}[WW^\top A(x)]$, where architectural differences across models are encapsulated by the structure matrix $A(x)$, and training dynamics are described by the order parameter $M = WW^\top$. By integrating symmetry analysis, differential-geometric expansions, and Lotka–Volterra dynamics, they derive—for the first time—a universal leading-order form grounded in smoothness and symmetry principles. The theory accurately predicts the temporal ordering of mode activation and the emergent power-law exponent after mode merging, with numerical validation across diverse architectures and training configurations.
📝 Abstract
Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training losses instead follow smooth power laws. Variants of both behaviors occur in architectures with very different microscopic structures, which is the signature of a few relevant collective variables. We show that a symmetry fixes what those variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero weights present at the start of training, the quadratic $\Tr[WW^{\top}A(x)]$, in which every architectural detail is confined to a single ``structure matrix" $A(x)$ that we compute for each architecture. Perceptrons, attention layers, mixtures of experts, and convolutions become one model at different $A$. Its training dynamics then close on the ``order parameter" $M=WW^{\top}$ and, whenever the data matrices share an eigenbasis, reduce to a Lotka--Volterra equation whose modes switch on one after another. The smaller the initial weights, the further apart the switch-on times, and the plateaus appear as a singular limit of a smooth flow; when many modes are unresolved the same events merge into a power law in training time whose exponent the theory predicts. We confirm both numerically across training methods and architectures.
Problem

Research questions and friction points this paper is trying to address.

sudden learning
scaling laws
neural networks
gradient descent
collective variables
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural Quadratic Forms
symmetry
scaling laws
Lotka–Volterra dynamics
structure matrix