SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析高维多类逻辑回归的训练动态,揭示了基于梯度优化下的交叉熵风险的精确缩放规律,展示了学习过程中的阶段性和模型容量与优化之间的关系。
📝 Abstract
We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws governing the cross-entropy risk under gradient-based optimization. We show that learning proceeds sequentially across classes, from most to least frequent. When the class priors follow a power law distribution, the risk dynamics decompose into three phases: an initial plateau until the first class is learned, a power-law decay regime during which sequential learning occurs, and a final convergence regime. We then analyze how model capacity interacts with optimization under a fixed compute budget. When the effective dimension is restricted via projection onto leading principal components, the risk decomposes into a capacity term (a power law in the retained dimension) and an optimization term (a power law in training time). Optimizing this tradeoff yields a compute-optimal scaling law for logistic regression, with explicit prescriptions for model size and training time as functions of compute. These results extend theoretical scaling laws from linear regression to multiclass classification, while connecting to empirical scaling laws observed in large-scale neural networks.
Problem

Research questions and friction points this paper is trying to address.

multiclass logistic regression
scaling laws
gradient-based optimization
cross-entropy risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

multiclass logistic regression
sequential learning
scaling laws
power law distribution
model capacity
🔎 Similar Papers
2024-10-02International Conference on Machine LearningCitations: 1