Towards a Theoretical Understanding to the Generalization of RLHF
This work investigates the generalization ability of Reinforcement Learning from Human Feedback (RLHF) in high-dimensional large language models. Departing from conventional analyses that rely on the consistency of maximum likelihood estimation, the study establishes, for the first time, a generalization bound under an end-to-end RLHF framework by leveraging the algorithmic stability perspective. By introducing a linear reward model and a feature coverage condition, and analyzing both Gradient Ascent (GA) and Stochastic Gradient Ascent (SGA) algorithms, the authors prove that when the feature coverage condition holds, the generalization error of the empirical optimal policy converges at a rate of $O(n^{-1/2})$. This result further extends to policies obtained via gradient-based optimization, thereby providing theoretical justification for the generalization performance of practical RLHF implementations.