To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究针对AI系统中难以优化的用户偏好问题,提出了一种名为CurriPO的方法,通过构建多样化的用户奖励模型课程来提高整体满意度并减少训练时间。
📝 Abstract
Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developing data-efficient and accurate reward models that capture minority preferences despite scarce data. We push this line of inquiry one step further and argue that data-efficient and accurate per-user reward models are not sufficient: users whose reward models are difficult to \textit{optimize} at the policy level can become a new underserved group. We start from the observation that one user's reward model can be easy to optimize from the initial policy while another's is not. We argue that, given a sufficiently diverse user population, a curriculum naturally emerges between easy- and hard-to-optimize reward models. Building on this insight, we propose CurriPO, which grows a tree-structured curriculum to accommodate diverse user-specific objectives, covering the population in a single traversal. Specifically, CurriPO automatically constructs a curriculum over diverse user reward models, allowing it to branch from the existing curriculum and reuse reward models previously incorporated into the curriculum. To the best of our knowledge, this is the first work to explicitly exploit multi-user structure to address optimization in AI alignment. Extensive experiments on personalized continuous control in a simulated environment show that CurriPO achieves $1.2$--$2.1\times$ the population satisfaction of the strongest baseline while substantially reducing training time. Additional analysis attributes much of this improvement to the users left underserved by conventional optimization.
Problem

Research questions and friction points this paper is trying to address.

reward model
policy optimization
user preferences
curriculum learning
AI alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Curriculum Learning
Diverse Preferences
Reward Optimization
AI Alignment
Policy Optimization
🔎 Similar Papers
No similar papers found.