World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain

๐Ÿ“… 2026-09-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡็ป“ๅˆ้ข„ๆต‹ไธ–็•Œๆจกๅž‹ๅ’ŒPPO็ญ–็•ฅ๏ผŒ่งฃๅ†ณไบ†ไบบ็ฑปๆœบๅ™จไบบๅœจ่„š่ธ็บฆๆŸๅœฐๅฝขไธŠ็š„่กŒ่ตฐ้—ฎ้ข˜๏ผŒๆ้ซ˜ไบ†่ทจ่ถŠ้šœ็ข็š„ๆˆๅŠŸ็އๅ’Œๆ•ˆ็އใ€‚
๐Ÿ“ Abstract
Foothold-constrained terrain is characterized by sparse, discontinuous, or geometrically restricted feasible foot contacts, as encountered on stepping stones, across gaps, and on narrow stair treads. On such terrain, a single misstep often leaves little room to recover, so policies that base foot-placement decisions primarily on the immediately visible terrain are prone to failure. We ask whether a learned predictive summary of near-future observations and rewards can provide the anticipatory information required in such settings. We present World-Model-Augmented Visual Locomotion (WM-LOCO), which jointly trains a recurrent world model and a PPO policy. Conditioned on proprioception and a single onboard depth image, the world model produces a predictive recurrent feature that guides the policy, without explicit foothold labels. In simulation, WM-LOCO succeeds on gaps and stepping stones where a matched baseline fails completely, and matches the baseline's success rate on stairs while improving stride efficiency and reducing pelvis acceleration. We deploy the same policy onboard a physical Unitree G1 humanoid using onboard proprioception and a single depth stream; it traverses all three terrain classes with an average success rate of 93.3%.
Problem

Research questions and friction points this paper is trying to address.

Foothold-Constrained Terrain
Visual Locomotion
Humanoids
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Model-Augmented
Visual Locomotion
Recurrent World Model
Foothold-Constrained Terrain
PPO Policy
๐Ÿ”Ž Similar Papers