DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决7-DoF末端执行器控制中细粒度动作预测不准确问题,提出DUET-DINO模型,通过同时从侧视和腕部相机视角学习动作条件预测,提高机器人操作中的潜在规划性能。
📝 Abstract
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO
Problem

Research questions and friction points this paper is trying to address.

action-conditioned latent world models
7-DoF end-effector control
fine-grained spatial and rotational actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Conditioning
Latent Planning
7-DoF Control
Action-Conditioned Predictions
Dual-Camera Setup
💼 Related Jobs
No related jobs found.
N
Nisarga Nilavadi
Artificial Intelligence and Robotics Lab, University of Technology Nuremberg (UTN), Germany.
Ralf Römer
Ralf Römer
Technical University of Munich
Machine LearningRoboticsEmbodied AIVLAControl
Moritz Reuss
Moritz Reuss
Karlsruhe Institute of Technology
Roboticsmachine learning
M
Michael Krawez
Artificial Intelligence and Robotics Lab, University of Technology Nuremberg (UTN), Germany.
T
Tobias Jülg
Artificial Intelligence and Robotics Lab, University of Technology Nuremberg (UTN), Germany.
A
Angela P. Schoellig
Learning Systems and Robotics Lab, Technical University of Munich (TUM), Germany.
Rudolf Lioutikov
Rudolf Lioutikov
TT-Professor, Intuitive Robots Lab, Karlsruhe Institute of Technology
Machine LearningRoboticsRobot LearningReinforcement LearningImitation Learning
Wolfram Burgard
Wolfram Burgard
Professor of Computer Science, University of Technology Nuremberg
RoboticsArtificial IntelligenceAIMachine LearningComputer Vision