DIDO: Distilling Interaction-Centric Dynamics into One-Step Denoising for World Action Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多步视频模型在机器人操作中引入额外延迟的问题,DIDO通过单步去噪浓缩互动中心动力学,同时保持场景结构和减少推理延迟。
📝 Abstract
World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-centric dynamics most critical for manipulation. To address this issue, we propose DIDO, which distills the converged dynamics of a multi-step video model into a single denoising step. DIDO combines distribution matching distillation with interaction-centric representation guidance. Beyond compressing multi-step generation into one forward pass, DIDO explicitly models the gripper, manipulated object, and their interaction using supervised bounding-box visual reasoning tokens. Additionally, DIDO aligns the target object's representations across multiple model layers with features from a pretrained DINOv3 encoder. This interaction-centric guidance helps the distilled model preserve both the relevant entities and their future dynamics in a single step, while substantially reducing inference latency. DIDO achieves an average success rate of 99.0\% on LIBERO, 76.6\% on LIBERO-Plus, and 92.0\% on RoboTwin, while also demonstrating effective transfer to long-horizon and generalization tasks in real-world robotic manipulation.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
denoising
interaction dynamics
latency
manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

DIDO
one-step denoising
interaction-centric dynamics
distribution matching distillation
DINOv3 encoder
🔎 Similar Papers
No similar papers found.
Jing Lyu
Jing Lyu
Shanghai Jiao Tong University
Power electronicsstabilityrenewable energy grid integrationhigh-voltage dc transmission
Shuanghao Bai
Shuanghao Bai
Xi'an Jiao Tong University Phd student
Vision Language ModelsDomain AdaptationDomain GeneralizationRobotic Manipulation
R
Runze Xiao
Beijing Academy of Artificial Intelligence (BAAI)
Zhenyu Liao
Zhenyu Liao
Applied Scientist in Amazon Inc.
optimizationmathematics
W
Wenxing Tan
Beijing Academy of Artificial Intelligence (BAAI)
Z
Zihan Tang
Tsinghua University
R
Ruochuan Shi
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Academy of Artificial Intelligence (BAAI)
C
Cheng Peng
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Beijing Academy of Artificial Intelligence (BAAI)
Yuheng Ji
Yuheng Ji
Institute of Automation, Chinese Academy of Sciences
Embodied AIComputer Vision
Y
Yihao Wang
Beijing Academy of Artificial Intelligence (BAAI)
Badong Chen
Badong Chen
Professor of Xi'an Jiaotong University, Xi'an, China
signal processingmachine learningbrain machine interfacesrobotics
Pengwei Wang
Pengwei Wang
University of Calgary
Computer Science Security
Zhongyuan Wang
Zhongyuan Wang
BAAI
Knowledge MiningDatabaseNLPText Understanding
Xiaoguang Zhao
Xiaoguang Zhao
Tsinghua University
MEMSMicrosystemsTHzMetamaterialWireless communication