Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

📅 2026-02-17
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
为了解决大语言模型难以适应多样化个人偏好的问题,本文提出了一种新的个性化群体相对策略优化框架P-GRPO,通过针对不同偏好组的历史奖励进行优势估计,从而更好地学习和对齐不同的用户偏好。
📝 Abstract
Despite their sophisticated general-purpose capabilities, Large Language Models (LLMs) often fail to align with diverse individual preferences because standard post-training methods, like Reinforcement Learning with Human Feedback (RLHF), optimize for a single, global objective. While Group Relative Policy Optimization (GRPO) is a widely adopted on-policy reinforcement learning framework, its group-based normalization implicitly assumes that all samples are exchangeable, inheriting this limitation in personalized settings. This assumption conflates distinct user reward distributions and systematically biases learning toward dominant preferences while suppressing minority signals. To address this, we introduce Personalized GRPO (P-GRPO), a novel alignment framework that decouples advantage estimation from immediate batch statistics. By normalizing advantages against preference-group-specific reward histories rather than the concurrent generation group, P-GRPO preserves the contrastive signal necessary for learning distinct preferences. We evaluate P-GRPO across diverse tasks and find that it consistently achieves faster convergence and higher rewards than standard GRPO, thereby enhancing its ability to recover and align with heterogeneous preference signals. Our results demonstrate that accounting for reward heterogeneity at the optimization level is essential for building models that faithfully align with diverse human preferences without sacrificing general capabilities.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Preference Alignment
Reinforcement Learning
Heterogeneous Preferences
Group Relative Policy Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Personalized GRPO
heterogeneous preferences
advantage estimation
reward heterogeneity
contrastive signal
🔎 Similar Papers
J
Jialu Wang
Apple Inc.
H
Heinrich Peters
Apple Inc.
A
Asad A. Butt
Apple Inc.
N
Navid Hashemi
Apple Inc.
A
Alireza Hashemi
Apple Inc.
P
Pouya M. Ghari
Apple Inc.
J
Joseph Hoover
Apple Inc.
James Rae
James Rae
Apple Inc.
M
Morteza Dehghani
Apple Inc.