FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of simultaneously achieving motion alignment, identity preservation, and visual realism in music-driven dance generation. We propose a parallel pose-RGB dual-stream diffusion framework that integrates timestep-aware pose injection with a persistent identity mechanism to enable joint modeling of 3D motion and 2D visuals. Additionally, we construct a high-resolution in-the-wild dance dataset. By effectively combining explicit motion control with reference image synthesis, our method demonstrates superior performance in both dance generation and video synthesis tasks. The proposed approach significantly enhances temporal coherence and identity consistency, ultimately facilitating high-fidelity music-driven dance video generation.
📝 Abstract
Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.
Problem

Research questions and friction points this paper is trying to address.

Music-driven dance video generation
Identity-preserving human animation
Temporal coherence
Music-to-motion correspondence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Pose and RGB Streams
Timestep-aware Pose Injection
Persistent Identity Injection
Music-driven Dance Video Generation
In-the-wild Dance Dataset
🔎 Similar Papers
G
Genying Li
School of Artificial Intelligence, Beijing University of Posts and Telecommunications; Beijing Academy of Artificial Intelligence
B
Boda Lin
School of Artificial Intelligence, Beijing University of Posts and Telecommunications
J
Jiachen Li
School of Artificial Intelligence, Beijing University of Posts and Telecommunications
Z
Zijian Jia
School of Artificial Intelligence, Beijing University of Posts and Telecommunications; Beijing Academy of Artificial Intelligence
H
Haojie Zheng
Beijing Academy of Artificial Intelligence; School of Software and Microelectronics, Peking University
Yiming Wang
Yiming Wang
School of Chemical Engineering, East China University of Science and Technology
lifelike soft materialsnon-equilibrium materialssupramolecular self-assembly
S
Shuchen Weng
Beijing Academy of Artificial Intelligence
S
Si Li
School of Artificial Intelligence, Beijing University of Posts and Telecommunications