WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in the evaluation of action-conditioned world models (ACWMs), which has predominantly emphasized visual fidelity or task performance while neglecting their core function as physical simulators—namely, the causal fidelity between actions and environmental responses. To this end, we formalize the notion of an "observable simulator contract" and introduce WorldSimProbe, a fine-grained diagnostic framework that assesses ACWMs across five dimensions: local control sensitivity, global trajectory variation, multi-source action consistency, interaction grounding, and dynamics. Built upon controlled testing protocols, our framework leverages calibration analysis, dense action-motion correspondence, and spurious interaction detection. Evaluated on RoboTwin, ManiSkill, and LIBERO across six open-source models (>18,000 instances), it reveals systematic deficiencies in action execution, interaction grounding, and dynamics, with results strongly aligned with human judgment and downstream task performance.
📝 Abstract
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.
Problem

Research questions and friction points this paper is trying to address.

simulator faithfulness
action-conditioned world models
embodied manipulation
physical simulation
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

WorldSimProbe
action-conditioned world models
simulator fidelity
Observable Simulator Contract
embodied manipulation
P
Peterson Co
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; EvoPhys AI; Joy Future Academy, JD
S
Sicheng Hu
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; EvoPhys AI
C
Chunxuan Jiao
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; EvoPhys AI
H
Hongyang Cheng
EvoPhys AI
Yulin Luo
Yulin Luo
Peking University
Data-centric AILLMVLMEmbodied AI
Yijie Xu
Yijie Xu
Hong Kong University of Science and Technology (Guangzhou)
Data MiningNatural Language ProcessingLarge Language Models
Sixiang Chen
Sixiang Chen
The Hong Kong University of Science and Technology (Guangzhou)
Computer VisionImage RestorationAIGCMLLM
Z
Zhongxia Zhao
EvoPhys AI
Zihao Wang
Zihao Wang
HKUST
Machine LearningLogical ReasoningOptimal Transport
D
DaFeng Chi
Joy Future Academy, JD
Peidong Liu
Peidong Liu
Westlake University
3D computer visionRobotics
Y
YuTong Chen
EvoPhys AI; Beijing Institute of Technology
H
Henghua Liu
EvoPhys AI; Beijing Institute of Technology
Zhihao Yuan
Zhihao Yuan
Ph.D student at The Chinese University of Hong Kong, Shenzhen
Vision and Language3D Scene Understanding
H
Huizhu Jia
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Yuzheng Zhuang
Yuzheng Zhuang
Senior Researcher @ Huawei Noah's Ark Lab
Reinforcement LearningOptimizationAutonomous DrivingCommunication
T
Tianle Zhang
Joy Future Academy, JD
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
Huajie Tan
Huajie Tan
Peking University
Embodied AIFoundation Models
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models