🤖 AI Summary
This work addresses the lack of a theoretical framework unifying the control of information propagation, gradient flow, and learning dynamics in deep neural network initialization. By integrating mean-field theory with random matrix theory in the infinite-width-and-depth limit, we establish—for the first time—a direct link between correlation propagation and the Neural Tangent Kernel (NTK), revealing that learning dynamics at criticality are governed by output correlations. We demonstrate the equivalence of information propagation and learning dynamics in the infinite-depth regime and prove that orthogonal initialization effectively suppresses dominant finite-size effects introduced by Gaussian initialization. These theoretical predictions are quantitatively validated in networks of finite width and depth, highlighting the pivotal role of orthogonal initialization in shaping the asymptotic learning dynamics of deep networks.
📝 Abstract
The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive. Combining mean-field theory and random matrix theory, we establish a direct link between correlation propagation and the Neural Tangent Kernel (NTK) that governs learning in the sequential limit of infinitely wide, infinitely deep networks. Correlation propagation to infinite depth is possible only at a single critical point in the weight-bias variance plane. At this point, we show that the end-to-end Jacobian vanishes algebraically with depth, and use this to prove that the NTK becomes exactly proportional to the output correlation at infinite depth. This equivalence between information propagation and learning dynamics had not yet been noticed. We further show that orthogonal initialisation suppresses the leading finite-size corrections present under Gaussian initialisation, clarifying the respective roles of the two initialisation ensembles in this limit. These theoretical predictions are validated quantitatively on finite-width, finite-depth networks. Together, these results demonstrate that orthogonal initialisation at criticality plays a central role in controlling the asymptotic dynamics of deep learning.