RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出RoboGesture框架,通过设计数据、建模和控制来解决人形机器人同步生成语义手势的问题,采用层次语义-声学对齐器和流式条件运动生成器等方法。
📝 Abstract
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the "modality eclipse" where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

Human-robot Interaction
Co-speech Gestures
Semantic Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Semantic-Acoustic Aligner
Streaming Conditional Motion Generator
Anti-Inertia CFG Masking
MPC-based safety filter
🔎 Similar Papers
No similar papers found.