HarnessWAM: Bridging Prediction and Deliberation in World Action Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of World Action Models (WAMs) in complex embodied tasks—specifically, their inadequate planning, state maintenance, and failure recovery stemming from a disconnect between prediction and reasoning. To overcome this, we propose an agent-based framework that employs a vision-language model–driven task manager to maintain a structured scene belief and task graph. High-level semantic plans are projected into sequences of atomic skills that respect both task dependencies and robot capability constraints. An event-driven dual-timescale feedback mechanism, coupled with a lightweight progress estimator, enables a verifiable and recoverable execution loop. Our approach introduces, for the first time, external structured state maintenance and closed-loop decision-making, substantially enhancing WAMs’ global planning and local fault tolerance. Experiments show our method achieves a 59.6% end-to-end success rate (69.9% on subtasks) on RoboMemArena and 23.7% on RoboCerebra Ideal, significantly outperforming existing approaches.
📝 Abstract
World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. However, finite-horizon prediction and action generation are insufficient for complex embodied tasks that require global planning, cross-stage state maintenance, execution verification, and failure recovery. We refer to this mismatch as the prediction-deliberation gap of WAMs. To address this gap, we propose HarnessWAM, an agentic framework for WAMs. HarnessWAM employs a vision-language-model-based Task Manager to maintain an evidence-grounded scene belief and a structured task graph. A capability-conditioned executable-space projection further constrains open-ended semantic plans into sequences of atomic skills that satisfy task dependencies, embodiment-state constraints, and the capability boundary of the underlying WAM. During execution, HarnessWAM operates through an event-driven, dual-timescale feedback loop: a lightweight progress estimator continuously provides high-frequency execution evidence, while the Task Manager deliberates at salient milestones by jointly considering the current observation, task state, and interaction history to determine whether to advance the task, acquire additional observations, revise the plan, or initiate local recovery. This mechanism enables the robot to recover its state after a subtask failure and resume execution without discarding previously acquired scene knowledge. HarnessWAM achieves state-of-the-art full-task and subtask success rates of 59.6% and 69.9% on RoboMemArena, and an SR of 23.7% on RoboCerebra Ideal. These results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
prediction-deliberation gap
embodied task execution
state maintenance
failure recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
agentic framework
structured task graph
dual-timescale feedback
capability-conditioned planning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhaopeng Gu
Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Bingke Zhu
Bingke Zhu
Institute of Automation,Chinese Academy of Science
T
Tianxi Lin
Yinwang Intelligent Technology Co., Ltd., Shenzhen, China; Beijing Institute of Technology, Beijing, China
Guibo Zhu
Guibo Zhu
Institute of Automation, Chinese Academy of Sciecnes
Artificial IntelligenceComputer VisionMachine Learning
Y
Yingying Chen
Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
K
Kai Wang
Institute of Automation, Chinese Academy of Sciences, Beijing, China; Yinwang Intelligent Technology Co., Ltd., Shenzhen, China
T
Tingyu Yuan
Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
Chaoyang Zhao
Chaoyang Zhao
Institute of Automation, Chinese Academy of Sciences
computer vision
Zhaowen Li
Zhaowen Li
National Laboratory of Pattern Recognition,Institute of Automation,Chinese Academy of Sciences
Computer VisionArtificial IntelligenceSelf-supervised Learning
Peng Su
Peng Su
Ph.D at The Chinese University of Hong Kong
Deep LearningPhysical AIAutonomous Driving
J
Jinqiao Wang
Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; Wuhan AI Research, Wuhan, China