DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过将视频扩散模型重新用于几何编码,解决了自我中心视角下3D手部运动恢复中的遮挡和视野外问题,提出DreamHand框架,显著提高了手部轨迹恢复的准确性。
📝 Abstract
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce DreamHand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. DreamHand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that needs no test-time camera intrinsics. Across five egocentric benchmarks, DreamHand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%-61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.
Problem

Research questions and friction points this paper is trying to address.

Egocentric Video
3D Hand Motion Recovery
Occlusion
Out-of-sight Gaps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deterministic Clean-Latent Encoder
Bidirectional Spatiotemporal Decoder
Video Diffusion Models
Egocentric 3D Hand Motion Recovery
Occlusion-Robust
Y
Yufei Liu
Shanghai Jiao Tong University
Xixi Wang
Xixi Wang
University of Rochester
Machine learningPattern analysisAlzheimer's disease
H
Hao Li
The Chinese University of Hong Kong
G
Ganlong Zhao
The Chinese University of Hong Kong
K
Kaitong Cai
ACE Robotics
C
Chengkai Jin
Nanyang Technological University
Chunxiao Liu
Chunxiao Liu
Senior Research Director, SenseTime
Autonomous DrivingPrediction/Decision/Planning/ControlReinforcement LearningAI Agent
Jianbo Liu
Jianbo Liu
University of Notre Dame
S
Siyuan Huang
ACE Robotics
Xingang Pan
Xingang Pan
Assistant Professor, MMLab@NTU, Nanyang Technological University
Computer VisionDeep LearningComputer Graphics
H
Hongsheng Li
The Chinese University of Hong Kong