DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DiffSynth-Music框架,通过音频条件的KV缓存适配器增强音乐合成控制力,解决文本和歌词对音乐时间、旋律等控制不足的问题。
📝 Abstract
Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable audio conditioning to a music synthesis backbone through layer-wise key-value injection. The three template models, Control, Prosody, and Reference, are initialized from the backbone diffusion transformer and trained with conditional flow matching. They support five control types: beats, vocals, accompaniment, prosody, and reference audio. A shared variational autoencoder maps conditioning waveforms into a common latent space, enabling their attention memories to be combined. With the template timestep fixed at the clean-data endpoint and other inputs held constant, each control cache is computed once and reused throughout sampling. Training pairs are derived from music recordings using beat extraction, source separation, vocal resynthesis, and reference-excerpt selection. Single-control evaluations on Mandarin and English songs demonstrate improved adherence across all five control types and better lyric fidelity under vocal conditioning relative to the backbone. Automatic music-quality and instruction-following scores remain broadly comparable to those of the evaluated base models, with metric-specific trade-offs. We release the three template models to support research and creative applications in controllable music generation.
Problem

Research questions and friction points this paper is trying to address.

Music Generation
Audio Conditioning
Control Types
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-conditioned KV-cache adapters
composable audio conditioning
layer-wise key-value injection
conditional flow matching
shared variational autoencoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhongjie Duan
Zhongjie Duan
East China Normal University
Image Synthesis
S
Shengchuan Gao
Shanghai Jiao Tong University
H
Hong Zhang
Alibaba Group
Yingda Chen
Yingda Chen
Alibaba Group, Microsoft