HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决通用机器人策略在复杂长时任务中成功率低的问题,提出HaWMPO方法,通过感知预测幻觉调整策略优化,提高训练效率和成功率。
📝 Abstract
Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but such training relies on costly physical interactions, suffers from low sample efficiency, and may introduce hardware and safety risks. World models offer a promising alternative by enabling policy optimization with imagined rollouts. However, long-horizon rollouts generated by world models often suffer from prediction hal- lucinations, producing biased state transitions that can mislead policy learning. To address this issue, we propose Hallucination-aware World Model-based Pol- icy Optimization (HaWMPO), a closed-loop reinforcement learning pipeline for VLA policy post-training with world models. Specifically, HaWMPO introduces an action-conditioned hallucination-aware model to estimate the reliability of gen- erated image sequences, and incorporates hallucination scores into group relative policy optimization through a Reward-Soft mechanism, suppressing unreliable ac- tion chunks during training. On the LIBERO benchmark, HaWMPO achieves the best average success rate, with gains of 15.0% over the base model and 2.8% over the strongest baseline; real-world experiments on a G1 robot further validate its effectiveness, raising the average success rate on two manipulation tasks from 67.5% to 80.0%.
Problem

Research questions and friction points this paper is trying to address.

Generalist Robot Policy
Long-horizon Scenarios
World Models
Prediction Hallucinations
Policy Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucination-aware
World Model
Policy Optimization
Visual-Language-Action (VLA)
💼 Related Jobs
No related jobs found.
Z
Zengjue Chen
Joy Future Academy, JD; School of Artificial Intelligence, Jilin University
Peidong Liu
Peidong Liu
Westlake University
3D computer visionRobotics
J
Jiawei Li
Joy Future Academy, JD
Q
Qi Wang
School of Artificial Intelligence, Jilin University