Learning from Active Human Involvement through Proxy Value Propagation
This work addresses the insufficient policy alignment and safety in human-in-the-loop reinforcement learning (HIL-RL) under settings with no explicit reward signals. We propose Proxy Value Propagation (PVP), a novel method that encodes real-time human interventions and demonstrations as high/low binary value labels—without requiring an external reward function—and propagates these values across state-action pairs via temporal-difference (TD) learning. PVP is modular and seamlessly integrates with mainstream RL algorithms (e.g., SAC, PPO), and includes a lightweight human-in-the-loop interface. Experiments demonstrate that PVP significantly improves both policy safety and fidelity to human intent across continuous and discrete control benchmarks. Notably, in a realistic autonomous driving scenario within *Grand Theft Auto V*, PVP achieves rapid convergence with minimal human intervention, validating its practical deployability in complex, reward-free environments.