Iterative Audio Separation with Mixture Consistency via MIMO Model Extension

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种通过扩展到多输入多输出配置来实现稳定有效的迭代音频分离框架,解决混合一致性问题,并在实验中展示了性能提升。
📝 Abstract
This paper proposes a general framework for stable and effective iterative audio separation with mixture consistency by extending source separation models to a multi-input multi-output (MIMO) configuration. In the field of audio separation, mixture consistency is an essential property for many applications that require accurate phase and timbral information of target sources. While iterative approaches such as diffusion models achieve perceptually superior results in speech enhancement or user-guided target source separation tasks, most existing methods focus on single-step separation with a single-input single-output (SISO) or single-input multi-output (SIMO) configuration through architectural improvements, since mixture-consistent audio separation is generally regarded as a regression problem that admits a unique solution. By extending these architectures to a MIMO configuration, we introduce iterative prediction without compromising the architectural advantages or the characteristics of mixture consistency. We conduct a comprehensive ablation study of combining the framework with discriminators and extending it to a generative model. Experimental results demonstrate significant performance improvements when applying the proposed framework to state-of-the-art separation models.
Problem

Research questions and friction points this paper is trying to address.

audio separation
mixture consistency
iterative prediction
MIMO configuration
source separation
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterative audio separation
mixture consistency
MIMO configuration
generative model
🔎 Similar Papers
No similar papers found.