SalsaAgent: A multimodal embodied language model for interactive dance generation

๐Ÿ“… 2026-05-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenges of insufficient nonverbal responsiveness, coordination, and synchrony in human-robot partnered dancing by proposing a multimodal embodied language model. The approach extends the vocabulary of a large language model with discrete motion tokens and pairwise relational tokens, while incorporating audio input to enable real-time full-body salsa motion generation in response to both a human leader and background music. Key innovations include skeleton-dynamics-based automatic text descriptions for semantic alignment of tokens and a two-stage โ€œtoken-to-diffusionโ€ generation pipeline. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in motion quality, musical synchronization, partner coordination, and dyadic spatial consistency, as validated through comprehensive subjective and objective evaluations.
๐Ÿ“ Abstract
Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generates expressive, full-body salsa dance motions in reaction to a human leader and against a contextual music backdrop. We formulate interaction as nonverbal motion token passing, extending the vocabulary of a large language model (LLM) to process discrete motion tokens, pairwise relation tokens, and audio. Our contributions include new tokens for full-body and motion relations, LLM fine-tuning using automatically derived text descriptions of skeleton dynamics for token grounding, and a two-stage token-to-diffusion pipeline. Subjective and objective evaluations demonstrate the effectiveness of our approach in terms of motion quality, music and partner coordination, and consistent two-person spatial behavior, with significant improvements over baselines.
Problem

Research questions and friction points this paper is trying to address.

interactive dance generation
embodied language model
nonverbal interaction
motion coordination
multimodal perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal embodied language model
interactive dance generation
motion tokenization
LLM fine-tuning
token-to-diffusion pipeline
๐Ÿ”Ž Similar Papers
No similar papers found.
P
Payam Jome Yazdian
Simon Fraser University, Burnaby, Canada
Z
Zoe Stanley
Simon Fraser University, Burnaby, Canada
A
Angelica Lim
Simon Fraser University, Burnaby, Canada