Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文分析了RLHF在用户偏好不同时效用失真的问题,通过奖励裁剪方法改善了因分布不匹配导致的效用失真,并提出了优化建议。
📝 Abstract
While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Gölz et al. (2025) demonstrated that the \textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $β$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a consequence of distribution mismatch between the distribution generating preference data ($μ$) and the KL reference policy ($π_{\mathrm{ref}}$). To this end, we establish tight upper and lower bounds on the distortion of RLHF across multiple regimes of the KL regularization strength. We show that in a representative regime, under the Bradley-Terry model, the distortion is $\tildeΘ(βB + β)$, where $B$ is an upper bound on the log density ratio between $μ$ and $π_{\mathrm{ref}}$. In particular, when there is no distribution mismatch (i.e., $μ= π_{\mathrm{ref}}$), RLHF achieves the optimal distortion of $O(β)$ up to a constant. Our results suggest that, to reasonably maximize average utility with RLHF, it is preferable to use on-policy sampled preference data or to fine-tune before RLHF on data from a source close to $μ$.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning from Human Feedback
distortion
user utility
heterogeneous preferences
Bradley-Terry temperature parameter
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning from Human Feedback (RLHF)
distribution mismatch
KL regularization
Bradley-Terry model
on-policy sampled preference data