Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究了标准初始化下多项式宽度两层网络学习正交多指数模型的动力学,通过改进的梯度流方法证明了增量学习的存在,并揭示了参数质量的竞争再分配。
📝 Abstract
Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study training dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Our theoretical analysis uses slightly modified gradient flow, while vanilla gradient descent empirically exhibits the same qualitative dynamics. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.
Problem

Research questions and friction points this paper is trying to address.

incremental learning
orthogonal multi-index models
standard initialization
neuron interactions
feature learning dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

incremental learning
competitive reallocation of parameter mass
symmetry-based finite-width approximation
🔎 Similar Papers
2024-10-02International Conference on Machine LearningCitations: 1