SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of zero-shot compositional generation combining music-driven dance with precise vocal expression by proposing a unified video diffusion framework. The method introduces a novel character-aware audio conditioning mechanism that models vocals as semantic roles, integrating hard compact routing, frame-level joint audio injection, and asymmetric supervised training to achieve decoupled singing-dancing capabilities and compositional reasoning. Experimental results demonstrate superior performance in strong beat alignment, reliable character switching, and high-fidelity lip synchronization. Notably, the proposed model achieves these advancements with significantly fewer parameters than existing baselines, effectively resolving the problem of compositional zero-shot singing-dancing video generation while maintaining computational efficiency and generation quality.
📝 Abstract
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.
Problem

Research questions and friction points this paper is trying to address.

Singing-and-Dancing Video Generation
Compositional Zero-Shot
Role-Aware Audio Conditioning
Music-Conditioned Body Motion
Vocal Articulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Zero-Shot
Role-Aware Audio Conditioning
Unified Video Diffusion
Asymmetric Supervision
Hard-Compact Routing
🔎 Similar Papers
No similar papers found.