Institution profile

Zhongke FiveAges (Hangzhou) Intelligent Technology Co., Ltd.

Industry researchasia · cn
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Aug 06, 2026

This work investigates whether existing action-conditioned world models can generalize to unseen robot morphologies beyond mere visual memorization. To this end, the authors introduce XEWorld, a cross-embodiment evaluation benchmark that establishes, for the first time, an isolated-embodiment assessment paradigm. This framework systematically evaluates zero-shot and few-shot visual rendering capabilities of models when confronted with novel robots that share physical consistency but differ in embodiment structure. Experiments reveal that current models struggle to map abstract joint actions into coherent visual trajectories, relying heavily on visual similarity rather than kinematic or dynamic consistency for generalization. Furthermore, few-shot adaptation often leads to catastrophic forgetting of previously seen embodiments. These findings underscore the critical need for architectural innovations that explicitly disentangle appearance from physical dynamics.

0 citationsRead paper

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

Jun 25, 2026

This work addresses the limitations of existing world-action models, which struggle with long-horizon manipulation tasks that depend on early observations and task progression due to their reliance on short-term history and near-future predictions. To overcome this, the authors propose a memory-augmented world-action model that jointly captures multi-scale historical context, local future dynamics, and global task progress to enable event-memory-guided video and action denoising. Key innovations include a similarity-based multi-memory bank mechanism, a task-progress supervision objective, and a memory read-write strategy integrating visual event extraction with identity and temporal embeddings. Experiments demonstrate substantial performance gains, improving success rates on RMBench from 28.4% to 69.8%, and achieving 91.5% and 80.0% success rates on stage-wise and full-task evaluations, respectively, in real-world Franka robot experiments.

0 citationsRead paper
Recent publications

Latest Papers

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Aug 06, 2026

This work investigates whether existing action-conditioned world models can generalize to unseen robot morphologies beyond mere visual memorization. To this end, the authors introduce XEWorld, a cross-embodiment evaluation benchmark that establishes, for the first time, an isolated-embodiment assessment paradigm. This framework systematically evaluates zero-shot and few-shot visual rendering capabilities of models when confronted with novel robots that share physical consistency but differ in embodiment structure. Experiments reveal that current models struggle to map abstract joint actions into coherent visual trajectories, relying heavily on visual similarity rather than kinematic or dynamic consistency for generalization. Furthermore, few-shot adaptation often leads to catastrophic forgetting of previously seen embodiments. These findings underscore the critical need for architectural innovations that explicitly disentangle appearance from physical dynamics.

0 citationsRead paper

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

Jun 25, 2026

This work addresses the limitations of existing world-action models, which struggle with long-horizon manipulation tasks that depend on early observations and task progression due to their reliance on short-term history and near-future predictions. To overcome this, the authors propose a memory-augmented world-action model that jointly captures multi-scale historical context, local future dynamics, and global task progress to enable event-memory-guided video and action denoising. Key innovations include a similarity-based multi-memory bank mechanism, a task-progress supervision objective, and a memory read-write strategy integrating visual event extraction with identity and temporal embeddings. Experiments demonstrate substantial performance gains, improving success rates on RMBench from 28.4% to 69.8%, and achieving 91.5% and 80.0% success rates on stage-wise and full-task evaluations, respectively, in real-world Franka robot experiments.

0 citationsRead paper