Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了异步P2P网络中模型参数不同无法直接平均的问题,通过知识蒸馏方法在预测空间而非参数空间达成共识,证明了其收敛性。
📝 Abstract
Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) exchanges soft predictions rather than weights and sidesteps this obstacle, yet convergence theory for fully decentralized, asynchronous peer-to-peer (P2P) KD is lacking. We provide one, relocating consensus from parameter space to function (output) space: a KD event is a geometric contraction operator in logit space on the peers' predictive distributions, which we analyse in the Hilbert space of predictions on a reference measure. Under standard smoothness/variance assumptions and two realizability assumptions, one bridging parameter SGD to the functional step and one controlling restricted task/KD alignment, the time-averaged functional stationarity and function-space disagreement converge at rate $O(1/(ηT))$ to an $O(η)+O(B_f^2)+O(ζ_f^2)$ neighbourhood. Here $B_f$ is the distance from the task optimum to the peers' reachable classes and $ζ_f$ measures persistent local-task heterogeneity. Across homogeneous, width-heterogeneous, and mixed-family networks of the experiments, KD contracts function disagreement by $40-61\times$, while isolated training does not. The sampled stationarity diagnostic has late transient exponents $0.99-1.90$ on the shared-skeleton main runs, and the four-point step-size sweep exhibits the predicted transient: neighbourhood tradeoff.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Distillation
Asynchronous P2P Gossip Learning
Decentralized Learning
Convergence Theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
Decentralized Learning
Asynchronous P2P
Function Space Convergence
Gossip Network
🔎 Similar Papers
L
Lucas Qingyang Fang
University of California, Santa Cruz, USA
T
Tiyao Liu
China University of Petroleum, China
J
Jinhao Jing
The Chinese University of Hong Kong, Shenzhen, China
Z
Zeji Li
City University of Hong Kong, Hong Kong SAR, China
K
Kaijie Chen
University of Virginia, USA
Harikrishna Kuttivelil
Harikrishna Kuttivelil
University of California, Santa Cruz
distributed aiedge intelligencefederated learningiotedge intelligence applications
Katia Obraczka
Katia Obraczka
Professor of Computer Engineering, UC Santa Cruz
Computer Networks