Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过在潜空间中使用轻量级卷积探针引导音高内容,解决了扩散模型音乐生成中的可控性问题,无需重新训练或修改基础模型。
📝 Abstract
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based systems comparatively underexplored. We present a lightweight method for steering the pitch content of audio produced by Stable Audio Open, a latent diffusion model for music synthesis. A small convolutional probe containing approximately 125k parameters is trained to decode frame-level pitch-class activations from the model's variational autoencoder latent space, using paired audio and MIDI data. At inference time, the frozen probe serves as a differentiable loss function: its gradient with respect to the denoising latent is used to nudge generation toward a user-specified pitch-class sequence, requiring no retraining or architectural modification of the base model. Across 27 evaluation trials spanning 9 text prompts and 3 target melodies, probe-guided generation increases melodic coherence by 2.4x over the unguided baseline (p < 1e-5, Wilcoxon signed-rank test), demonstrating that musically meaningful structure is both recoverable and steerable in diffusion-based music latent spaces.
Problem

Research questions and friction points this paper is trying to address.

Diffusion-based Music Generation
Pitch-class Steering
Latent-space Probes
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent diffusion model
convolutional probe
pitch-class steering
differentiable loss function
melodic coherence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yushi Ye
Carnegie Mellon University
W
Wilson Zheng
Carnegie Mellon University
Yongyi Zang
Yongyi Zang
Smule, Inc.
Computer AuditionSpeech ProcessingMusic Information RetrievalMusic Composition