Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对强凸随机优化中固定时间保证与实际自适应停止决策之间的不匹配问题,通过构建可观察的、轨迹自适应的置信序列来解决,允许在保证精度的同时自适应停止SGD。
📝 Abstract
Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory. This mismatch creates a fundamental certification problem: fixed-time guarantees do not generally remain valid at data-dependent stopping times, while deterministic horizons derived from worst-case bounds can be highly conservative. We address this problem for strongly convex stochastic optimization by constructing fully observable, trajectory-adaptive upper confidence sequences for the squared distance of the last iterate to the optimizer and the suboptimality of a weighted average. These bounds hold simultaneously over time, attain the optimal $1/t$ decay rate up to iterated-logarithmic factors in the worst case, and adapt to the realized stochastic gradients, allowing SGD to stop as soon as a prescribed accuracy is certified without sacrificing statistical validity. Our approach treats the evolving SGD trajectory as a sequential experiment whose observations provide evidence about the unknown optimization error. To formalize this perspective, we develop new recursive confidence-sequence techniques and a general time-uniform empirical Bernstein inequality for adapted processes with time-varying conditional means and predictable ranges that may grow without bound. We further extend these confidence-sequence constructions to minibatch SGD, with the empirical Bernstein bounds exploiting the realized second-moment structure within each minibatch. Numerical experiments show that the resulting stopping rules can require several orders of magnitude fewer iterations than natural deterministic horizons.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Optimization
Adaptive Stopping
Confidence Sequences
SGD
Strongly Convex
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory-adaptive stopping rules
upper confidence sequences
time-uniform empirical Bernstein inequality
stochastic gradient descent (SGD)