MVWeaver: A Hierarchical Music Video Generation Agent with a Learned Song-to-Visual Bridge

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
MVWeaver通过结合分层规划与学习到的歌曲到视觉桥接方法,解决了自动生成音乐视频中的长形式连贯性和基于歌曲的视觉发展问题。
📝 Abstract
Music videos are an important form of audiovisual expression in contemporary culture. They translate and extend the expressive content of songs through deliberate visual design. Existing automatic music video (MV) generation systems can generate visually plausible shots, yet often struggle with long-form coherence and song-grounded visual development. We present MVWeaver, a music video generation agent that integrates hierarchical planning with a learned song-to-visual bridge that translates song understanding into executable shot plans. The MVWeaver architecture comprises a comprehensive song analysis module, a visual planner that constructs hierarchical plans, and downstream image and video generation models that render the planned content. To equip a general-purpose LLM with MV-specific song-to-visual knowledge, we learn a bridge between song analysis and visual planning from real-MV-derived supervision and curate 1,861 real-world song--MV pairs with structured song-side, MV-side, and teacher-inferred song-to-visual rationale annotations. Using these annotations, we perform LoRA-based supervised fine-tuning (SFT) of a large language model to predict song-to-visual bridges that guide hierarchical visual planning. Our experiments demonstrate stronger song-grounded visual translation, richer visual development, and greater conceptual and shot-to-shot coherence, while ablations support the benefits of learned bridge conditioning.
Problem

Research questions and friction points this paper is trying to address.

Music Video Generation
Long-form Coherence
Song-grounded Visual Development
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Planning
Song-to-Visual Bridge
LoRA-based Supervised Fine-Tuning
Comprehensive Song Analysis
💼 Related Jobs
No related jobs found.
S
Sifei Li
Institute of Automation, Chinese Academy of Sciences, China, School of Artificial Intelligence, University of Chinese Academy of Sciences, China, and KlingAI Research, China
Minyan Luo
Minyan Luo
MAIS, Institute of Automation, Chinese Academy of Sciences
Computer VisionAIGC
X
Xu Li
KlingAI Research, China
G
Guodong Qi
KlingAI Research, China
X
Xincan Wang
Shanghai Theatre Academy, China
Hanwen Wang
Hanwen Wang
Johns Hopkins University, SOM
Quantitative Systems PharmacologyOncologySystems Biology
C
Chen Zhang
KlingAI Research, China
Pengfei Wan
Pengfei Wan
Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
Oliver Deussen
Oliver Deussen
Professor of Computer Science, University of Konstanz
Computer GraphicsVisualizationModelling
W
Weiming Dong
Institute of Automation, Chinese Academy of Sciences, China and School of Artificial Intelligence, University of Chinese Academy of Sciences, China