Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过语音印象引导的伪三元组构建方法解决了方向跟随TTS中因缺少相对修改训练数据的问题,实现稳定且保持说话人身份的语音修改。
📝 Abstract
Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/
Problem

Research questions and friction points this paper is trying to address.

direction-following TTS
training data
relative modifications
Innovation

Methods, ideas, or system contributions that make the work stand out.

pseudo-triplet construction
direction-following TTS
impression-controllable TTS model
large language model (LLM)
🔎 Similar Papers
No similar papers found.