🤖 AI Summary
This study addresses the limitations of limited voice diversity and inadequate editing capabilities in text-to-speech generation by proposing a unified generative framework. By constructing a hybrid dataset comprising both real and synthetic voices and designing an improved diffusion Transformer architecture, this work achieves unified multi-task speech synthesis and editing. Experimental results demonstrate that the proposed method outperforms state-of-the-art models in prompt alignment, perceptual quality, and usability. Notably, it significantly enhances zero-shot voice cloning and flexible editing performance, thereby establishing a new paradigm for controllable speech synthesis.
📝 Abstract
Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, existing systems face two key challenges. First, they struggle to generate a diverse range of voices, spanning real-world human speakers and fictional characters. Second, they lack robust and flexible voice editing capabilities, such as voice cloning and the ability to modify attributes like emotion and tone. In this paper, we propose VoiceDesigner, a unified framework for text-to-voice generation and editing that supports diverse and controllable voice design. To tackle the above challenges, we propose solutions from two perspectives. First, we develop a hybrid data pipeline that leverages digital signal processing techniques and speech generation models to construct a diverse voice dataset covering both real-world and fictional voices. Second, we introduce a diffusion transformer with architectural improvements to better handle complex conditioning and enhance multi-task performance, enabling unified voice generation and editing. Through subjective and objective evaluations, VoiceDesigner achieves superior prompt alignment with both voice descriptions and editing instructions, while maintaining competitive perceptual quality and voice usability compared to state-of-the-art TTV models.