A Better Spur Should Start From Each Objective

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多目标强化学习中的稀疏奖励、奖励冲突等问题,提出MMPO框架,通过数据、梯度和约束层面的干预提高训练稳定性和性能。
📝 Abstract
Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing traditional linear scalarization to experience severe metric oscillations. To address optimization conflicts among multiple objectives in real-world deployment scenarios, we propose Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization. Specifically, MMPO performs exposure debiasing to mitigate sparse and biased rewards, applies priority-aware orthogonal projection to decouple conflicting gradients, and introduces self-prompted gradient constraints to prevent dominant objectives from overwhelming weaker ones. Experiments on real-world e-commerce datasets show that MMPO improves training stability and consistently achieves better performance across conflicting metrics. Moreover, it generalizes robustly to broader tasks such as ToolRL and code generation, demonstrating its effectiveness as a practical paradigm for multi-objective alignment.
Problem

Research questions and friction points this paper is trying to address.

Multi-Objective Reinforcement Learning
Sparse Rewards
Reward Conflicts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Marginal Preference Optimization
Exposure Debiasing
Priority-Aware Orthogonal Projection
Self-Prompted Gradient Constraints
🔎 Similar Papers
2024-04-122024 IEEE Intelligent Vehicles Symposium (IV)Citations: 8
S
Shanwen Mao
Harbin Institute of Technology, Harbin, China
H
Hao Zhang
Harbin Institute of Technology, Harbin, China
Guangtao Nie
Guangtao Nie
JD.com Inc., Beijing, China
Z
Zhiheng Li
Institute of Automation, Chinese Academy of Sciences, Beijing, China
H
Huimu Wang
Institute of Automation, Chinese Academy of Sciences, Beijing, China
Sulong Xu
Sulong Xu
京东
S
Simiu Gu
JD.com Inc., Beijing, China