World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出World-Coherent-Decoding方法,通过自验证测试时规划来提高机器人控制中视觉未来生成的可靠性,从而提升任务成功率。
📝 Abstract
World Action Models (WAMs) aim to control robots by stochastically generating visual futures and then decoding actions, but empirical observations indicate that the results can strongly depend on which future is selected. We propose World-Coherent-Decoding (WCD), a self-verifying test-time planning framework that treats WAM rollouts as falsifiable future--action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using internal generative signals: flow-based video surprisal for visual plausibility and action path effort for action-generation stability. After execution, the realized observation audits the selected imagination, yielding an imagination--reality mismatch that trains a lightweight online predictor for future candidate selection. Thus, WCD converts delayed self-verification into pre-execution reliability estimation without updating the backbone model. On RoboTwin 2.0, WCD improves Hard success under limited randomized-scene supervision from $55.80\%$ to $60.90\%$, with a $+16.43$ gains on Horizon-3 tasks, and shows qualitative robustness on real Franka visual-shift tests. These results highlight a simple principle: test-time scaling for WAMs depends less on sampling more futures than on selecting reliable ones.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Reliability
Future Selection
Test-time Planning
Robot Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Coherent-Decoding
self-verifying test-time planning
world action models
flow-based video surprisal
action path effort