How Edge of Stability Hinders SCAFFOLD in Federated Optimization

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在联邦优化中,边缘稳定性(EoS)如何影响SCAFFOLD算法估计全局梯度的能力,解释了其实际表现不如FedAvg的原因。
📝 Abstract
In federated learning, it is well known that heterogeneous data can (in theory) slow down optimization, and much effort has been directed at designing optimization algorithms that are unaffected by data heterogeneity, such as the SCAFFOLD algorithm. Yet, despite strong theoretical guarantees, SCAFFOLD does not usually outperform the much simpler FedAvg in practice. In this work, we propose that this gap is due to the presence of Edge of Stability (EoS) and progressive sharpening in federated optimization, supported by extensive empirical probing. First, we find that EoS-like dynamics occur with both FedAvg and SCAFFOLD under a variety of architectures and hyperparameters. We observe that the equilibrium value of the sharpness is inversely proportional to the learning rate (as in GD), and interestingly, the degree of data heterogeneity (but not the number of local steps) also affects the equilibrium value. Most importantly, we observe that SCAFFOLD's ability to estimate the gradient of the global objective is severely degraded at the EoS, as measured by the correlation between sharpness and SCAFFOLD's error in estimating the global gradient along the optimization trajectory. This suggests a mechanism for SCAFFOLD's lackluster performance in deep learning: with high sharpness at the EoS, SCAFFOLD cannot reliably estimate the global gradient.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Data Heterogeneity
Edge of Stability
SCAFFOLD
Innovation

Methods, ideas, or system contributions that make the work stand out.

Edge of Stability
Federated Learning
SCAFFOLD
Gradient Estimation
Data Heterogeneity
💼 Related Jobs
No related jobs found.
A
Anant Khandelwal
Georgia Institute of Technology, College of Computing
Michael Crawshaw
Michael Crawshaw
George Mason University
Machine learningoptimizationdeep learningfederated learning
M
Mingrui Liu
George Mason University, Department of Computer Science