Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Puppeteer模型,通过结合语音信号、动作历史、初始姿势和物体几何信息,在因果潜在空间中生成与物体相关的、符合姿势约束的同步手势。
📝 Abstract
Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.
Problem

Research questions and friction points this paper is trying to address.

co-speech gestures
temporal coherence
object-grounded
Innovation

Methods, ideas, or system contributions that make the work stand out.

Posture-Aware
Object-Grounded
Causal Latent Space
Conditional Diffusion
SceneGes
🔎 Similar Papers
No similar papers found.