JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions
This work addresses the challenges of high-fidelity music generation and cross-scene narrative coherence in long-form video scoring by proposing the JenBridge framework. The approach adopts a two-stage paradigm: it first pretrains on large-scale text–audio corpora to learn musical priors, then fine-tunes under dual text–visual conditioning to achieve cross-modal alignment. A key innovation is an adaptive transition mechanism driven by a large language model agent, which dynamically selects diverse transition strategies to handle scene shifts. Experimental results on a Transformer-based flow-matching generative model and the newly constructed LVS evaluation benchmark demonstrate that the proposed method significantly outperforms existing techniques in both objective metrics and subjective assessments, particularly excelling in transition smoothness and overall narrative coherence.