Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
This work addresses the challenge of scaling reinforcement learning to complex loco-manipulation tasks, where conventional approaches rely heavily on handcrafted dense rewards. The authors propose a novel framework that leverages sample-based model predictive control (SMPC) as an automated expert policy generator to efficiently construct large-scale offline datasets in simulation. This dataset is then used within a hierarchical architecture combining offline-to-online reinforcement learning under sparse rewards with a low-level dynamically stable controller. Notably, the method eliminates the need for manual reward engineering and enables agents trained solely with sparse rewards to surpass the performance of the SMPC teacher policy. The approach demonstrates strong empirical results, successfully deploying on both the Spot quadruped and G1 humanoid robots with high performance, robustness, and effective sim-to-real transfer.