Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Encore,通过自适应信号路由解决长时间同步音视频生成问题,增强局部连续性和全局一致性。
📝 Abstract
Existing audio-video generation methods produce well-synchronized clips but are limited to short durations, while long-video generation methods extend duration through chunk-based iterative synthesis yet lack audio entirely. Generating long audio-video jointly is fundamentally harder than either task alone: each chunk must simultaneously maintain video temporal coherence, audio temporal coherence, and cross-modal synchronization, whose conditioning signals enter the model through different pathways. In this work, we present Encore for long-form synchronized audio-video generation. Our key insight is to factor this challenge into: (1) local continuity handled by iterative generation with explicit cross-chunk context propagation, and (2) global consistency enforced via reference audio-video signals with shifted position embedding. Building on this design, we propose Adaptive Signal Routing (ASR), which introduces learnable attention biases within self-attention and learnable residual scales on cross-attention outputs, enabling the model to adaptively modulate the influence of each conditioning signal. Trained end-to-end for joint audio-video generation, Encore also supports infinite-length audio-to-video and video-to-audio synthesis at inference by conditioning on the ground-truth modality throughout the denoising process. Experiments on our extended VerseBench for long audio-video evaluation demonstrate that Encore significantly outperforms existing methods in both generation quality and temporal coherence. Code and data for this paper are at https://github.com/shaohua-pan/Encore.
Problem

Research questions and friction points this paper is trying to address.

audio-video generation
temporal coherence
cross-modal synchronization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Signal Routing (ASR)
long-form synchronized audio-video generation
cross-chunk context propagation
shifted position embedding
🔎 Similar Papers
No similar papers found.
S
Shaohua Pan
Baidu, China
J
Junbao Chen
Beijing Institute of Technology, China and Baidu, China
Shengyi He
Shengyi He
Columbia University
J
Jingfeng Xue
Beijing Institute of Technology, China
Wen Tao
Wen Tao
DATA61,CSIRO
Transport network modelling
Haocheng Feng
Haocheng Feng
Baidu
computer vision
S
Siming Fan
Baidu, China
D
Dongwei Pan
Baidu, China
Yi Yang
Yi Yang
Baidu Inc., Beijing, China
Large MultiModal LearningLarge Graph LearningVideo Retrieval
Wei He
Wei He
Baidu
Natural Language Processing
H
Hang Zhou
Baidu, China