Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling
This work addresses the challenge that existing singing voice conversion (SVC) methods struggle to reliably extract clean vocal melodies from accompanied recordings due to harmonic interference. To overcome this limitation, we propose a zero-shot, cross-lingual SVC system that explicitly models both the main melody and residual harmonics—a first in SVC—enabling effective processing of polyphonic audio. The architecture integrates a CQT-based pitch extractor, a stochastic sampler, and a conditional flow-matching diffusion decoder, jointly optimizing pitch, linguistic content, and time–frequency features. Experimental results demonstrate that our approach consistently outperforms current baselines on both harmonically rich and monophonic datasets, achieving superior performance in terms of naturalness, timbre similarity, and harmonic reconstruction fidelity.