Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过WorldEcho诊断行动条件世界模型在非专家行动下的表现,并提出WorldSync方法增强行动跟随,提高策略学习的可靠性。
📝 Abstract
Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated. To address this gap, we introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. Our diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories, either ignoring the commanded actions or producing visually invalid rollouts. We further propose WorldSync, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alignment. It broadens the training distribution over action consequences, grounds intermediate video representations in action-induced robot dynamics through an Action-Forcing Expert, and aligns predicted changes under action interventions with the corresponding changes in ground-truth futures. Experiments on RoboTwin benchmarks and real-robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to achieve higher success rates.
Problem

Research questions and friction points this paper is trying to address.

Action-Conditioned World Models
Off-Expert Actions
Policy Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

WorldEcho
WorldSync
Action-Conditioned Generation
Policy Learning
SE(3) Trajectory Alignment
🔎 Similar Papers
No similar papers found.
Sixiang Chen
Sixiang Chen
The Hong Kong University of Science and Technology (Guangzhou)
Computer VisionImage RestorationAIGCMLLM
J
Jiaming Liu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
J
Jixian Wu
Beijing Innovation Center of Humanoid Robotics
Yichen Guo
Yichen Guo
Master student in Nanyang Technological University
T
Tinghao Wang
Beijing Innovation Center of Humanoid Robotics
S
Siyuan Qian
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
H
Hao Chen
The Chinese University of Hong Kong
J
Jiajun Cao
Beijing Innovation Center of Humanoid Robotics
J
Jian Tang
Beijing Innovation Center of Humanoid Robotics
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models