UniVoice: A Unified Model for Speech and Singing Voice Generation
Traditional approaches struggle to jointly model natural speech and controllable singing within a single framework due to the fundamentally different constraints imposed by prosody and melody. This work proposes UniVoice, a unified generative architecture based on conditional flow matching that decomposes conditioning signals into content, melody, and timbre, and employs a shared DiT backbone for multimodal fusion. A key innovation is the introduction of learnable null melody tokens to replace explicit melody inputs during speech synthesis, thereby preserving explicit melodic control for singing while avoiding redundant constraints on speech; theoretical analysis shows this design approximates marginalization over melody. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26%, comparable to specialized TTS systems, and a singing PER of 16.22%, substantially outperforming the unified baseline Vevo1.5 (24.72%).