Deep Reinforcement Learning for Bipedal Locomotion: A Brief Survey
Existing deep reinforcement learning (DRL) frameworks for bipedal robot multi-task motion control suffer from poor generalization, low sim-to-real transfer efficiency, and a lack of systematic benchmarking. Method: We propose the first unified DRL taxonomy tailored to bipedal control, rigorously delineating the trade-offs between end-to-end and hierarchical control across task coverage, policy interpretability, and sim-to-real transfer. Our integrated framework synergizes PPO/SAC, hierarchical RL, model predictive control (MPC), and neural policy representation, accompanied by deployment principles balancing robustness and scalability. Contribution/Results: Leveraging a structured evaluation matrix spanning 30+ state-of-the-art methods, we identify multi-task generalization and cross-domain transfer as the two fundamental bottlenecks. This work establishes both theoretical foundations and engineering guidelines for DRL-driven embodied intelligent control.