DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DreamX-Creator 1.0系统,通过联合生成音频和视频流解决视觉动态与声学事件建模不互惠的问题。
📝 Abstract
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
Problem

Research questions and friction points this paper is trying to address.

audio-video generation
synchronized audio and video
2K resolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gated Cross-Modal Attention
Unified Audio-Video Data System
Progressive Joint Training
Audio-Video Reinforcement Learning
Autoregressive 1-Step 2K Refinement
🔎 Similar Papers
No similar papers found.
J
Jiashu Zhu
DreamX Team, Alibaba Group
Y
Yanhao Zheng
DreamX Team, Alibaba Group
R
Ruitian Tian
DreamX Team, Alibaba Group
R
Rujing Dang
DreamX Team, Alibaba Group
Shen Zhang
Shen Zhang
MEGVII
Deep LearningComputer Vision
B
Bingze Song
DreamX Team, Alibaba Group
J
Jiachen Lei
DreamX Team, Alibaba Group
R
Ruimin Lin
DreamX Team, Alibaba Group
J
Jiahong Wu
DreamX Team, Alibaba Group
X
Xiangxiang Chu
DreamX Team, Alibaba Group