Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization with Optimal Noise Scheduling
This work addresses the poorly understood implicit smoothing mechanism of stochastic gradient descent (SGD) in non-convex optimization. We provide the first theoretical characterization of how the SGD noise—determined jointly by learning rate, batch size, and gradient variance—induces a quantifiable smoothing effect on the objective function, and establish intrinsic links among smoothness, sharpness, and generalization performance. We propose a progressive optimization algorithm featuring joint scheduling of learning rate and batch size to dynamically modulate noise intensity for optimal implicit regularization. Empirical validation on ResNet-based image classification demonstrates significantly improved convergence stability and test accuracy; moreover, smoothness strongly correlates with generalization accuracy. Our core contributions are threefold: (i) an explicit, interpretable smoothing interpretation of SGD noise; (ii) a sharpness-driven theoretical framework for generalization; and (iii) the first noise-adaptive optimization paradigm with rigorous theoretical justification.