Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
This work addresses the tendency of existing deep reinforcement learning–based navigation methods to prioritize task objectives at the expense of social compliance in dense crowds, often resulting in behaviors that violate human social norms. To mitigate this issue, the authors propose a differentiable reward modeling approach grounded in Hall’s proxemic theory, formalizing personal space as a radial Gaussian mixture field. This formulation enables the computation of a local social cost within the robot’s field of view, which is seamlessly integrated into a deep reinforcement learning framework. The method uniquely translates proxemic theory into a dense, interpretable, and differentiable reward signal, significantly improving social compliance across diverse crowd densities and environments while maintaining navigation efficiency comparable to state-of-the-art approaches.