🤖 AI Summary
This work addresses a critical gap in the evaluation of action-conditioned world models (ACWMs), which has predominantly emphasized visual fidelity or task performance while neglecting their core function as physical simulators—namely, the causal fidelity between actions and environmental responses. To this end, we formalize the notion of an "observable simulator contract" and introduce WorldSimProbe, a fine-grained diagnostic framework that assesses ACWMs across five dimensions: local control sensitivity, global trajectory variation, multi-source action consistency, interaction grounding, and dynamics. Built upon controlled testing protocols, our framework leverages calibration analysis, dense action-motion correspondence, and spurious interaction detection. Evaluated on RoboTwin, ManiSkill, and LIBERO across six open-source models (>18,000 instances), it reveals systematic deficiencies in action execution, interaction grounding, and dynamics, with results strongly aligned with human judgment and downstream task performance.
📝 Abstract
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remains difficult to establish because prevailing evaluations emphasize visual quality, task outcomes, or coarse rollout-level responsiveness without directly testing simulator fidelity. To address this gap, we evaluate ACWMs through the observable capabilities expected of physical simulators. Accordingly, we formalize Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion. To operationalize this contract, we introduce WorldSimProbe, comprising five controlled suites spanning local control sensitivity, global trajectory variation, source-diverse actions, interaction grounding, and dynamics. Suite-specific evaluators assess simulator-relative calibration, dense action-to-motion correspondence, false-interaction grounding, and primitive-level dynamics. We evaluate six open-source ACWMs on more than 18,000 instances across RoboTwin, ManiSkill, and LIBERO. World-SimProbe reveals systematic action-realization degradation across control variation, structured failures in interaction grounding and dynamics, and benchmark signals consistent with human judgments and downstream outcomes. Together, this capability-based framework provides a transparent, and standardized paradigm for diagnosing ACWM simulator fidelity beyond coarse, task-directed evaluation.