Sparse In-Network Learning via Shortest-Path Backpropagation and Finite-Rate Gating
This work addresses the high communication overhead in distributed neural modular training by proposing a communication-efficient training framework. It constructs a capacity-aware shortest-path tree rooted at a fusion node and prunes non-tree edges to yield a sparse communication topology. Local routing is modeled via a finite-rate stochastic gating mechanism, and rate–distortion theory guides the joint optimization of sparsification and information compression. The method reduces training communication volume by 70.4% without compromising model accuracy and further decreases the transmission rate of latent variables by 45.7% through information bottleneck regularization, substantially enhancing the efficiency of distributed training.