EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了TTS系统中情感过渡不自然的问题,通过多通道混合管道、双阶段VAD条件和方向-幅度解耦注入方法提高了语音合成中的情绪转换质量。
📝 Abstract
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-pass flow blending pipeline synthesizes frame-aligned transition audio, circumventing the scarcity of natural intra-utterance transitions; (2) dual-stage Valence-Arousal-Dominance (VAD) conditioning guides prosodic planning in the LLM and acoustic realization in the flow decoder via frame-level VAD embeddings; (3) direction-magnitude decoupled injection structurally separates emotion direction from injection magnitude, preventing content degradation. EmoTra-TTS adds only +0.43% parameters with no latency overhead, achieves 30%-87% relative improvement on emotion transition quality, corroborated by 64.4%-79.5% overall win rates in pairwise preference tests against four SOTA baselines and two commercial systems.
Problem

Research questions and friction points this paper is trying to address.

Emotional TTS
Intra-utterance Emotion Transitions
Continuous Affect
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-pass flow blending
dual-stage VAD conditioning
direction-magnitude decoupled injection
🔎 Similar Papers
No similar papers found.