CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CAST方法,通过结合规划引导行为和当前策略来改进值学习,以解决模型基础强化学习中的动作选择问题。
📝 Abstract
Model-based reinforcement learning (MBRL) is a family of RL methods that learn a model of the environment and use it for action selection, making it well suited to robotics due to its sample efficiency. Combining learned models with online planning can further improve action selection, as the planner can exploit the model to find better actions than the learned policy alone. Recent methods combining learned policies with online planning typically learn the value of the policy rather than the stronger planner-guided behavior. We present CAST (Critic with Alternating State-value Target), which uses planner-guided behavior to improve value learning while regularizing the value estimate with the current policy. CAST replaces the action-value critic with a state-value critic, trained using a target that combines a real planner-guided transition and an imagined transition under the current policy. The resulting value function corresponds to an alternating process between planner-guided behavior and the current policy, allowing it to benefit from the stronger planner behavior while being regularised by the policy being learned. We evaluate CAST on the DeepMind Control and HumanoidBench Suites against several state-of-the-art methods, and demonstrate successful transfer to a physical Unitree Go2 quadruped performing a dynamic handstand.
Problem

Research questions and friction points this paper is trying to address.

model-based reinforcement learning
online planning
value learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

planner-guided behavior
state-value critic
alternating process
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pietro Noah Crestaz
Industrial Engineering Department, University of Trento, Trento, Italy; LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France
M
Mohamed Yassine Kabouri
LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France; Machines in Motion Laboratory, New York University, New York, USA
Nicolas Mansard
Nicolas Mansard
LAAS-CNRS, ANITI
Robotics
Andrea Del Prete
Andrea Del Prete
Associate Professor, University of Trento
Robotics