Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of selecting the latent space dimension in randomized low-dimensional reparameterization to ensure neural networks can be efficiently trained into low-loss regions. By characterizing accessibility phase transitions through conic geometry, the authors propose a directionally resolved quadratic theoretical framework that accurately predicts residual errors in random slices. Integrating structured random projections—such as Hadamard or recycled Gaussian mappings—with matrix-free curvature approximations and optimizer state compression, they develop a memory-efficient training framework. The method automatically determines the optimal dimensionality without exhaustive scanning, and empirical results on both vision and language models reveal training phase transitions that align closely with theoretical predictions, substantially outperforming existing approximation strategies that neglect directional information.
📝 Abstract
Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped into a full parameter update by a frozen random map. This raises a practical question: how large must the latent search space be to reach a low-loss region? We first express the known accessibility transition in an equivalent conic form, centered for compact convex targets at the statistical dimension of the polar cone. Our main theoretical contribution is an orientation-resolved quadratic master formula that predicts the random-slice residual from both the curvature spectrum and the reference-to-solution displacement profile. It yields a self-consistent isotropic-orientation predictor and, in a conservative radius-only specialization, recovers the earlier Gaussian-width quadratic bound. Building on this analysis, we introduce Random Mapping Networks (RaMaN), which instantiate the predicted latent dimension using structured Hadamard or seed-regenerated Gaussian maps. These constructions avoid the O(dP) storage of dense random maps and reduce optimizer-state memory from O(P) to O(d). We also develop matrix-free curvature approximations and sweep-free dimension selection. Across controlled quadratic and neural-curvature experiments, the orientation-resolved predictor closely tracks measured transition locations and outperforms orientation-agnostic approximations when displacement direction matters. End-to-end experiments further show sharp, protocol-dependent training transitions across image and language models.
Problem

Research questions and friction points this paper is trying to address.

random reparameterization
latent dimension
neural network training
low-loss region
accessibility transition
Innovation

Methods, ideas, or system contributions that make the work stand out.

random reparameterization
orientation-resolved prediction
structured random maps
latent dimension selection
matrix-free curvature approximation
💼 Related Jobs
No related jobs found.
A
Andrew Cheng
Department of Computer Science, Tsinghua University, Beijing, China; Department of Computer Science, The University of Manchester, Manchester, UK
A
Ali Eslamian
Department of Computer Science, University of Kentucky, Lexington, KY, USA
Jie Cheng
Jie Cheng
Institute of Automation, Chinese Academy of Sciences
Reinforcement Learning
M
Mehdi Zargham
Department of Computer Science, University of Dayton, Dayton, OH, USA
Qiang Cheng
Qiang Cheng
U Kentucky
data miningmachine learningpattern recognitionsignal/image processingbiomedical informatics