🤖 AI Summary
In high-dimensional learning, model stability undergoes a sharp phase transition when the sample size $n$ falls below a critical threshold—caused by the weakest Fisher direction being overwhelmed by sampling noise, rendering parameters unidentifiable. To address this, we develop the first non-asymptotic, necessary, and verifiable stability criterion, grounded in the minimal Fisher eigenvalue, and establish a finite-sample phase transition theory that precisely characterizes the phase boundary at the $d/n$ scale. We further propose Fisher-floor regularization—a novel, smoothness- and preprocessing-invariant spectral robustness diagnostic. Integrating Fisher information spectrum analysis, finite-sample random matrix theory, and high-dimensional statistical inference, we empirically validate our framework on Gaussian mixture and logistic regression models: the predicted phase-transition threshold cleanly separates reliable estimation from instability collapse, markedly enhancing interpretability and reliability in high-dimensional modeling.
📝 Abstract
In high-dimensional learning, models remain stable until they collapse abruptly once the sample size falls below a critical level. This instability is not algorithm-specific but a geometric mechanism: when the weakest Fisher eigendirection falls beneath sample-level fluctuations, identifiability fails. Our Fisher Threshold Theorem formalizes this by proving that stability requires the minimal Fisher eigenvalue to exceed an explicit $O(sqrt{d/n})$ bound. Unlike prior asymptotic or model-specific criteria, this threshold is finite-sample and necessary, marking a sharp phase transition between reliable concentration and inevitable failure. To make the principle constructive, we introduce the Fisher floor, a verifiable spectral regularization robust to smoothing and preconditioning. Synthetic experiments on Gaussian mixtures and logistic models confirm the predicted transition, consistent with $d/n$ scaling. Statistically, the threshold sharpens classical eigenvalue conditions into a non-asymptotic law; learning-theoretically, it defines a spectral sample-complexity frontier, bridging theory with diagnostics for robust high-dimensional inference.