CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text-to-speech systems struggle to achieve fine-grained prosodic control at the word or phoneme level while preserving the target speaker’s vocal timbre. To address this challenge, this work proposes CtrlSpeech, a framework built upon the DiTAR architecture that jointly models a global speaker embedding with phoneme-aligned multidimensional prosodic features—specifically pitch, loudness, and duration—to enable expressive speech synthesis from coarse to fine granularity. Notably, CtrlSpeech is the first approach to support phoneme-level local prosody controllability in zero-shot voice cloning, significantly enhancing the flexibility and practicality of expressive control without compromising naturalness. Experimental results validate the effectiveness of the proposed method.
📝 Abstract
Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.
Problem

Research questions and friction points this paper is trying to address.

expressive speech synthesis
fine-grained control
prosodic control
text-to-speech
zero-shot TTS
Innovation

Methods, ideas, or system contributions that make the work stand out.

coarse-to-fine control
expressive speech synthesis
prosodic control
zero-shot TTS
phone-aligned prosody
🔎 Similar Papers
No similar papers found.