Gradient Descent with Stochastic Subspaces via Persistence of Memory

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入'持久记忆'技术改进随机子空间方法,利用弱相关向量指导梯度下降过程,有效解决大规模优化问题。
📝 Abstract
Stochastic subspace methods have gained popularity as gradient descent based techniques for large scale optimisation problems, especially in distributed settings. In this paper, we introduce the technique of "persistence of memory" to greatly extend and improve the random subspace methods. To this end, we leverage a vector that is only weakly correlated with the gradient in order to provide a guiding structure to the generative process of the random subspace along which the descent is going to take place. This guidance vector may be fixed for a large number of iterations, only to be refreshed at wide intervals (on whose size we can provide guarantees in terms of problem parameters). In important machine learning settings, such as optimisation problems embodying sparsity or a minibatch structure, we show that the guidance vector can be obtained in an effective and computationally inexpensive manner by leveraging the structured properties of the problem. En route, we establish to our knowledge the first theoretical analysis of classical SSD methods for sparse functions. In a local neighbourhood of the optimum, we demonstrate an alignment phenomenon of our gradient estimates with a low-lying eigenvector of the Hessian, allowing a once-for-all computation of the guidance vector which renders the method computationally favourable even in scenarios with unstructured objectives.
Problem

Research questions and friction points this paper is trying to address.

Gradient Descent
Stochastic Subspaces
Large Scale Optimisation
Distributed Settings
Persistence of Memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

persistence of memory
stochastic subspace methods
guidance vector
sparse functions
alignment phenomenon
🔎 Similar Papers
S
Subhroshekhar Ghosh
Department of Mathematics, National University of Singapore, Singapore
C
Clement Z. Q. Ng
Department of Mathematics, National University of Singapore, Singapore
Pierre-Louis Poirion
Pierre-Louis Poirion
RIKEN
Mathematical OptimizationO.R.
A
Akiko Takeda
Center for Advanced Intelligence Project, RIKEN, Tokyo, Japan; Department of Mathematical Informatics, The University of Tokyo, Tokyo, Japan