🤖 AI Summary
This work addresses the limitations of existing convergence analyses for Muon optimizers in non-convex optimization, which often rely on strong assumptions and yield coarse bounds. We propose a streamlined analytical framework that dispenses with restrictive conditions on the update rule and directly characterizes convergence behavior under general non-convex settings. Our approach significantly tightens the upper bound on the convergence rate and applies to a broader class of non-convex optimization problems. Furthermore, it delivers a unified and refined theoretical guarantee for first-order methods based on orthogonalization, enhancing both the generality and precision of convergence analysis in this domain.
📝 Abstract
The Muon optimizer has recently attracted attention due to its orthogonalized first-order updates, and a deeper theoretical understanding of its convergence behavior is essential for guiding practical applications; however, existing convergence guarantees are either coarse or obtained under restrictive analytical settings. In this work, we establish sharper convergence guarantees for the Muon optimizer through a direct and simplified analysis that does not rely on restrictive assumptions on the update rule. Our results improve upon existing bounds by achieving faster convergence rates while covering a broader class of problem settings. These findings provide a more accurate theoretical characterization of Muon and offer insights applicable to a broader class of orthogonalized first-order methods.