🤖 AI Summary
This work addresses the scarcity of real-world human-centric data and the challenge of sim-to-real transfer in human-robot object handover prediction. To this end, the authors introduce Hand2Bot, the first RGB-D handover video dataset that captures both full-body poses and facial expressions. They further propose PassGen, a generative framework that integrates a stable video diffusion model with an intention-aware temporal face encoder, augmented by a morphological depth-editing strategy to emulate realistic depth sensor noise. This approach generates handover sequences exhibiting consistent hand-object interactions and authentic sensor artifacts. Evaluated on a physical robot platform, PassGen enables zero-shot transfer, achieving earlier intention anticipation, higher recognition accuracy, and significantly lower false-trigger rates compared to conventional hand-centric baselines.
📝 Abstract
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.