FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning

πŸ“… 2024-10-28
πŸ›οΈ ACM Multimedia
πŸ“ˆ Citations: 37
✨ Influential: 2
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the high communication overhead of large-scale models, such as Vision Transformers, in federated learning by proposing Federated Block Coordinate Gradient Descent (FedBCGD) and its accelerated variant, FedBCGD+. The method introduces, for the first time in federated learning, a block-wise parameter communication mechanism that uploads only a subset of parameter blocks per round, combined with stochastic variance reduction and client drift control strategies. Theoretical analysis shows that the communication complexity is reduced by a factor of 1/N compared to existing methods, where N denotes the number of blocks. Experimental results demonstrate that the proposed algorithms achieve faster convergence and higher communication efficiency than current state-of-the-art approaches.

Technology Category

Application Category

πŸ“ Abstract
Although Federated Learning has been widely studied in recent years, there are still high overhead expenses in each communication round for large-scale models such as Vision Transformer. To lower the communication complexity, we propose a novel Federated Block Coordinate Gradient Descent (FedBCGD) method for communication efficiency. The proposed method splits model parameters into several blocks including a shared block and enables uploading a specific parameter block by each client, which can significantly reduce communication overhead. Moreover, we also develop an accelerated FedBCGD algorithm (called FedBCGD+) with client drift control and stochastic variance reduction. To the best of our knowledge, this paper is the first work on parameter block communication for training large-scale deep models. We also provide the convergence analysis for the proposed algorithms. Our theoretical results show that the communication complexities of our algorithms are a factor 1 /N lower than those of existing methods, where N is the number of parameter blocks, and they enjoy much faster convergence than their counterparts. Empirical results indicate the superiority of the proposed algorithms compared to state-of-the-art algorithms.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
Communication Efficiency
Large-scale Models
Communication Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
Block Coordinate Descent
Communication Efficiency
Variance Reduction
Client Drift Control
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
J
Junkang Liu
School of Artificial Intelligence, Xi’an, Xidian University, China
Fanhua Shang
Fanhua Shang
Professor at Tianjin University
Machine LearningData MiningComputer Vision
Y
Yuanyuan Liu
School of Artificial Intelligence, Xi’an, Xidian University, China
Hongying Liu
Hongying Liu
Tianjin University
Machine learningImage processing
Y
Yuangang Li
University of Southern California, Los Angeles, US
Y
YunXiang Gong
School of Artificial Intelligence, Xi’an, Xidian University, China