🤖 AI Summary
To address the dual challenges of environmental noise interference and inadequate expressiveness modeling in real-world singing voice conversion (SVC), this paper proposes R2-SVC—a robust, expressive SVC framework. First, it constructs a noise-robust training dataset by simulating realistic degradations—including music separation artifacts and random fundamental frequency perturbations—and filtering samples via DNSMOS. Second, it integrates domain-adapted speaker representation learning with a neural source-filter (NSF) architecture to explicitly disentangle harmonic and noise components, thereby enhancing timbral controllability and naturalness. Evaluated on a multi-noise-condition SVC benchmark, R2-SVC achieves state-of-the-art performance, significantly improving conversion quality, inference stability, and cross-condition generalization. Notably, it is the first work to systematically bridge the gap between clean-data training and noisy real-world inference, advancing practical SVC deployment.
📝 Abstract
In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.