DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the rigid interactions in existing conversational synthesis caused by reliance on manual rules. We propose a decoupled framework separating content, temporal dynamics, and acoustics. By leveraging Large Language Models for script generation and full-duplex models to enable natural emergence of dialogue timing independent of text, we achieve high-fidelity TTS re-rendering alongside a newly constructed doctor-patient corpus. This approach pioneers dynamic temporal generation based on full-duplex interaction, effectively eliminating the mechanical artifacts inherent in traditional concatenation methods. Experimental results demonstrate that the generated dialogues significantly outperform baselines in both dynamic interactivity and naturalness, validating the efficacy of our proposed framework in producing more authentic conversational speech.
📝 Abstract
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, making conversational timing prescribed rather than interaction-driven. We present DuplexGen, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics. An LLM first generates the dialogue script, and then two full-duplex conversational models perform the script while listening to each other in real time. This allows conversational timing to emerge naturally while preserving the scripted content. Finally, a high-fidelity text-to-speech model re-renders the interaction without altering its timing. As a demonstration of the proposed framework, we construct a patient--clinician conversational speech corpus with construction-time annotations, including word timestamps, speaker activity, overlap regions, and interaction events. Experimental results show that the proposed framework produces conversational dynamics closer to real dialogue than conventional stitching-based synthesis.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Dialogue Speech
Conversational Timing
Interaction-driven
Full-duplex Conversation
Innovation

Methods, ideas, or system contributions that make the work stand out.

DuplexGen
Full-duplex conversational models
Decoupling content timing acoustics
Synthetic dialogue speech
Interaction-driven timing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.