Geometry of learning dynamics: Gradient descent versus natural gradient on the ridge of optimization

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过几何分析对比梯度下降与自然梯度下降方法在高容量联想记忆模型优化过程中的表现,揭示了自然梯度下降能更稳定、高效地解决学习动力学问题。
📝 Abstract
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit a "Ridge of Optimization" characterized by extreme stability and a highly skewed weight spectrum. However, the dynamical process by which learning converges to this critical regime has remained unclear. This paper provides a geometric analysis of the learning trajectories on the statistical manifold of a KLR-trained Hopfield network. By comparing the paths of Gradient Descent (GD) and Natural Gradient Descent (NGD), we elucidate the mechanisms governing the optimization process. Our analysis reveals that learning on the Ridge proceeds in two distinct phases. We show that the extreme curvature of the Ridge causes standard GD to follow a highly oscillatory, non-geodesic path. In stark contrast, NGD explicitly corrects for this geometry, following the ideal geodesic path and completely overcoming the instabilities faced by GD. We demonstrate experimentally that NGD not only converges significantly faster but also achieves a solution with superior generalization performance. These results establish that the highly structured geometry of the Ridge is optimally suited for information-geometric optimization, providing a new perspective on the interplay between learning dynamics and emergent representation geometry.
Problem

Research questions and friction points this paper is trying to address.

Gradient Descent
Natural Gradient Descent
Optimization Ridge
KLR
Hopfield Network
Innovation

Methods, ideas, or system contributions that make the work stand out.

Natural Gradient Descent
Optimization Ridge
Geodesic Path
Learning Dynamics
Generalization Performance