Institution profile

Interdisciplinary Program in Artificial Intelligence

Academic institution
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction

May 29, 2025

This work introduces MGE-LDM—the first unified latent diffusion framework addressing the fragmentation among music generation, source completion, and query-driven source separation. Methodologically, it reformulates all three tasks as conditional inpainting in the latent space, enabled by multi-condition text guidance, cross-dataset heterogeneous alignment, and joint modeling of mixtures, submixes, and isolated sources—achieving fully instrument-agnostic, end-to-end training. Trained jointly on Slakh2100, MUSDB18, and MoisesDB, MGE-LDM supports zero-shot, text-driven separation of arbitrary instruments, controllable mixture generation, and missing-source completion. Experiments demonstrate substantial improvements in audio fidelity and functional consistency over staged baselines. To our knowledge, this is the first framework realizing, within a single model and architecture, instrument-agnostic integration of these three core music signal processing tasks.

0 citationsRead paper
Recent publications

Latest Papers

MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction

May 29, 2025

This work introduces MGE-LDM—the first unified latent diffusion framework addressing the fragmentation among music generation, source completion, and query-driven source separation. Methodologically, it reformulates all three tasks as conditional inpainting in the latent space, enabled by multi-condition text guidance, cross-dataset heterogeneous alignment, and joint modeling of mixtures, submixes, and isolated sources—achieving fully instrument-agnostic, end-to-end training. Trained jointly on Slakh2100, MUSDB18, and MoisesDB, MGE-LDM supports zero-shot, text-driven separation of arbitrary instruments, controllable mixture generation, and missing-source completion. Experiments demonstrate substantial improvements in audio fidelity and functional consistency over staged baselines. To our knowledge, this is the first framework realizing, within a single model and architecture, instrument-agnostic integration of these three core music signal processing tasks.

0 citationsRead paper