A Survey of Reinforcement Learning from Human Feedback
Reinforcement learning (RL) often relies on hand-crafted reward functions that struggle to align with complex, nuanced human values. Method: This work systematically reviews RL from Human Feedback (RLHF), integrating reinforcement learning, Bayesian inference, preference modeling, reward modeling, and human-in-the-loop evaluation to support heterogeneous, multi-source feedback. It introduces the first unified, cross-task and cross-modal analytical framework for RLHF—extending beyond traditional preference-based RL (PbRL) limitations. Contribution/Results: The framework establishes a rigorous theoretical foundation and practical roadmap for human-AI value alignment. It clarifies the technical evolution, identifies core challenges (e.g., feedback sparsity, bias propagation, scalability), and proposes a standardized taxonomy for RLHF research. This serves as a comprehensive guide for algorithm design, ethical assessment, and real-world deployment—enabling principled, scalable, and value-aligned AI systems.