Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过Dynin-Omni模型解决了视觉目标和动态预测问题,使用统一的扩散视觉-语言-动作模型进行动作生成和选择。
📝 Abstract
Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluation, and joint refinement of action and future-state predictions. We continually pretrain the model on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets and adapt it separately to downstream domains. On two VLABench tasks, robot pretraining improves adaptation within a fixed Stage-2 step budget, and the full objective mixture improves shifted-instruction success over Policy-only post-training under the same coupled decoder. Combining goal guidance with joint action-next-state denoising further improves shifted-instruction success over action-only decoding; the benefit depends on how the predictions are composed. Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot. An optimized block-parallel implementation accelerates model-side action decoding by up to 29.2x relative to the base implementation under the reported profiling setup. These results support shared trajectory modeling as a common interface for learning complementary robot objectives and composing their predictions during control.
Problem

Research questions and friction points this paper is trying to address.

Visual Goal Prediction
Dynamics Prediction
Language-Conditioned Policies
Shared Trajectory Model
Omnimodal Diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omnimodal Unified Diffusion
Shared Trajectory Model
Action Prediction
Next-observation Prediction
Goal-state Prediction
🔎 Similar Papers
2024-04-02IEEE/RJS International Conference on Intelligent RObots and SystemsCitations: 0
💼 Related Jobs
No related jobs found.
H
Hoeun Lee
AIDAS Lab, Seoul National University
J
Jaeik Kim
AIDAS Lab, Seoul National University
J
Jusang Oh
AIDAS Lab, Seoul National University
J
Jinhyeok Kim
AIDAS Lab, Seoul National University
G
Geon Choi
AIDAS Lab, Seoul National University
H
Hyeonggeun Kim
AIDAS Lab, Seoul National University
Jaeyoung Do
Jaeyoung Do
Department of Electrical and Computer Engineering, Seoul National University
Generative AI (LLMs)Multi-Modal AI (NLP/Vision)Big Data Systems