🤖 AI Summary
This study addresses the challenge of simultaneously achieving motion alignment, identity preservation, and visual realism in music-driven dance generation. We propose a parallel pose-RGB dual-stream diffusion framework that integrates timestep-aware pose injection with a persistent identity mechanism to enable joint modeling of 3D motion and 2D visuals. Additionally, we construct a high-resolution in-the-wild dance dataset. By effectively combining explicit motion control with reference image synthesis, our method demonstrates superior performance in both dance generation and video synthesis tasks. The proposed approach significantly enhances temporal coherence and identity consistency, ultimately facilitating high-fidelity music-driven dance video generation.
📝 Abstract
Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.