VIBE: Video Instruction-aligned Background music gEneration

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视频转音乐模型缺乏语义控制和指令违规问题,VIBE通过动态跨层调节机制和综合奖励模型优化生成背景音乐。
📝 Abstract
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.
Problem

Research questions and friction points this paper is trying to address.

video-to-music
semantic control
instruction violations
reconstruction objectives
cross-modal conditioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditioning Connection
reward modeling taxonomy
text-and-video-to-music generation