Federated Compositional Muon Optimizer for Matrix-Wise Models
This work addresses the challenge of combinatorial optimization for matrix-valued models in distributed settings under non-independent and identically distributed (non-IID) data and non-convex objectives. The authors propose the Federated Combinatorial Muon optimizer (FedCoMuon) and its variance-reduced variant, FedCoMuon-VR, which uniquely integrate combinatorial gradient tracking and orthogonal momentum mechanisms into federated learning, augmented with momentum-based variance reduction. Theoretical analysis demonstrates that FedCoMuon-VR achieves a sample complexity of $O(\varepsilon^{-3})$ under non-IID non-convex conditions, improving upon the existing FedMuon method. Empirical evaluations on robust federated learning and task-distributed risk-sensitive meta-learning benchmarks show substantial gains over current combinatorial optimization baselines, establishing state-of-the-art accuracy.