KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决对话中重叠语音合成问题,提出KABURI-TTS方法,通过基于音素和语音活动条件的双通道发音渲染,实现更自然的人类对话模拟。
📝 Abstract
Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps that occur while the interlocutor is speaking. In this work, aiming at conversational speech synthesis that reproduces human-like overlap, we propose KABURI-TTS. KABURI-TTS takes a per-speaker phoneme raster as input and renders the speech of the two speakers on separate channels, conditioned on the per-frame phonemes and the voice activity derived from them. Because the phoneme raster is supplied by a separate module, the proposed method enables controllable generation of one-speaker-per-channel, two-party spoken dialogue. A user evaluation shows that, compared with strong baselines, the proposed method attains higher naturalness at both the utterance and the interaction level. Furthermore, an analysis of voice activity confirms that the proposed method produces more overlap and more frequent turn-taking.
Problem

Research questions and friction points this paper is trying to address.

conversational speech synthesis
full-duplex spoken dialogue
voice activity
phoneme raster
overlaps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phoneme-Keyed
Bi-channel Utterance Rendering
Activity-conditioned
🔎 Similar Papers
No similar papers found.