XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
This work investigates whether existing action-conditioned world models can generalize to unseen robot morphologies beyond mere visual memorization. To this end, the authors introduce XEWorld, a cross-embodiment evaluation benchmark that establishes, for the first time, an isolated-embodiment assessment paradigm. This framework systematically evaluates zero-shot and few-shot visual rendering capabilities of models when confronted with novel robots that share physical consistency but differ in embodiment structure. Experiments reveal that current models struggle to map abstract joint actions into coherent visual trajectories, relying heavily on visual similarity rather than kinematic or dynamic consistency for generalization. Furthermore, few-shot adaptation often leads to catastrophic forgetting of previously seen embodiments. These findings underscore the critical need for architectural innovations that explicitly disentangle appearance from physical dynamics.