MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

๐Ÿ“… 2026-08-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the inefficiency and task interference in multi-task learning that arise from neglecting the structural properties of model parameter matrices. It introduces matrix geometry into multi-objective optimization for the first time, proposing an orthogonally normalized gradient update mechanism grounded in matrix-valued steepest descent theory under the spectralโ€“nuclear norm geometry. The method guarantees convergence to Pareto stationary points even in non-convex settings. Empirical evaluations across multiple benchmarks demonstrate substantial improvements in both optimization efficiency and multi-task performance. Theoretically, the approach achieves convergence rates of $\mathcal{O}(T^{-1/2})$ in the deterministic setting and $\mathcal{O}(T^{-1/4})$ under stochastic gradients.
๐Ÿ“ Abstract
Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. Drawing from the theory of steepest descent for matrix-valued parameters, we propose MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. Theoretically, for smooth non-convex objectives, we establish convergence of the averaged Pareto-stationarity measure at rates of $\mathcal{O}(T^{-1/2})$ in the deterministic setting and $\mathcal{O}(T^{-1/4})$ under stochastic gradients. Empirical results across various benchmarks show that MOON consistently improves both optimization efficiency and final multi-task performance. Our code is available at https://github.com/KunlinLyu/MOON.
Problem

Research questions and friction points this paper is trying to address.

multi-objective optimization
multitask learning
matrix geometry
gradient manipulation
steepest descent
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-objective optimization
matrix geometry
orthonormalized gradients
multitask learning
steepest descent
๐Ÿ’ผ Related Jobs
No related jobs found.
Shiji Zhou
Shiji Zhou
Associate Professor, Beihang University
Online LearningStochastic OptimizationMulti-Objective OptimizationMulti-task Learning
K
Kunlin Lyu
Center for Applied Statistics, School of Statistics, Renmin University of China
Lei Zhang
Lei Zhang
Institute of Artificial Intelligence, Beihang University
R
Ruodong Wang
Center for Applied Statistics, School of Statistics, Renmin University of China
Y
Yifan Sun
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; Center for Applied Statistics, School of Statistics, Renmin University of China