Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过后训练框架和自然语言控制,解决细粒度情感和时长控制问题,使用监督微调和强化学习提高TTS模型的控制精度。
📝 Abstract
Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.
Problem

Research questions and friction points this paper is trying to address.

TTS
fine-grained control
emotion and duration
Innovation

Methods, ideas, or system contributions that make the work stand out.

post-training
fine-grained control
natural language
reinforcement learning
emotion and duration
🔎 Similar Papers
No similar papers found.