Movie Gen: A Cast of Media Foundation Models

📅 2024-10-17

🏛️ arXiv.org

📈 Citations: 65

✨ Influential: 9

career value

202K/year

🤖 AI Summary

This work addresses the challenge of high-quality multimodal video generation and editing by proposing a unified multimodal foundation model architecture. Methodologically, it introduces variable-aspect-ratio 1080p video latent-space modeling, cross-modal alignment training across text, image, video, and audio modalities, efficient tokenization, large-scale parallel training and inference optimization, and a rigorously quality-controlled data curation strategy coupled with a novel evaluation protocol. Key contributions include the first 30-billion-parameter video generation model supporting long-horizon generation (73K tokens, i.e., 16 seconds at 16 fps), instruction-driven precise editing, user-provided image personalization, and synchronized audio-video synthesis. The model achieves state-of-the-art performance across five benchmarks: text-to-video, video personalization, video editing, video-to-audio, and text-to-audio—demonstrating substantial improvements in temporal coherence and semantic controllability.

Technology Category

Application Category

📝 Abstract

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and generation of personalized videos based on a user's image. Our models set a new state-of-the-art on multiple tasks: text-to-video synthesis, video personalization, video editing, video-to-audio generation, and text-to-audio generation. Our largest video generation model is a 30B parameter transformer trained with a maximum context length of 73K video tokens, corresponding to a generated video of 16 seconds at 16 frames-per-second. We show multiple technical innovations and simplifications on the architecture, latent spaces, training objectives and recipes, data curation, evaluation protocols, parallelization techniques, and inference optimizations that allow us to reap the benefits of scaling pre-training data, model size, and training compute for training large scale media generation models. We hope this paper helps the research community to accelerate progress and innovation in media generation models. All videos from this paper are available at https://go.fb.me/MovieGenResearchVideos.

Problem

Research questions and friction points this paper is trying to address.

Develop high-quality 1080p HD video generation.

Enable precise instruction-based video editing.

Generate personalized videos using user images.

Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates 1080p HD videos

30B parameter transformer model

Precise instruction-based video editing

🔎 Similar Papers

MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies

2024-03-03arXiv.orgCitations: 21

Chrono: A Simple Blueprint for Representing Time in MLLMs

2024-06-26Citations: 4