CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes CookVoice, a unified non-autoregressive framework that enables multi-modal and multi-task human voice generation—including text-to-speech, text-to-singing, voice cloning, conversion, and editing—within a single model, addressing the limitations of existing systems that are often task-specific, autoregressive, and lacking in fine-grained control and inference efficiency. By decomposing voice into content, prosody, and style components and incorporating a frame-level alignment mechanism with an ordinary differential equation (ODE) solver, CookVoice achieves flexible control signal mapping and high-quality synthesis. With only 43.51 million parameters, the model generates both speech and singing in as few as four ODE steps, significantly enhancing controllability, generalization, and inference speed.
📝 Abstract
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, control signals, or autoregressive decoding, limiting fine-grained controllability and inference efficiency. In this paper, we propose CookVoice, a unified framework for multimodal, multi-style, and multi-task human voice generation. CookVoice decomposes the human voice into three key factors: content, prosody, and style, enabling both speech and singing voice generation within a unified model. To achieve precise and flexible controllability, we design a flexible alignment strategy that maps text, style, and prosody control signals onto the frame-level of spectrogram. This design allows CookVoice to support a wide range of tasks, including text-to-speech, text-to-singing voice, style-controllable generation, voice mimicry, voice conversion, and voice editing. Experimental results show that CookVoice achieves generation quality comparable to existing Text-to-Speech and text-to-singing voice baselines, while providing stronger style and prosody controllability. Moreover, CookVoice achieves comparable performance to large-scale baselines with only 43.51 million parameters and efficient inference using as few as 4 ODE steps, making it a practical solution for real-world human voice generation applications. Demo page is available at https://haoweilou.github.io/CookVoice/.
Problem

Research questions and friction points this paper is trying to address.

voice generation
controllability
unified framework
multi-modal
inference efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified framework
style controllability
prosody modeling
non-autoregressive generation
multimodal voice synthesis
🔎 Similar Papers
No similar papers found.