Bridging Data, Reasoning, and Alignment: A Unified Framework for Context-Aware Instruction-Following TTS

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决TTS系统在生成上下文相关语音时的局限性,提出了一种包含数据处理、上下文感知直接偏好优化及评估方法的系统优化方案。
📝 Abstract
The ISCSLP 2026 CoT-TTS Challenge requires TTS systems to generate Chain-of-Thought (CoT) reasoning from dialogue history before synthesizing contextually appropriate speech. While the official baseline establishes a unified architecture, it remains constrained by limited contextual comprehension, weak instruction fidelity, and suboptimal audio quality. We present a systematic optimization pipeline to address these limitations. First, we develop a data process framework that cleans raw data via FullSubNet denoising, Qwen3-ASR re-transcription, and Qwen3.5-35B-A3B-based history-CoT consistency analysis, while distilling 545K high-fidelity instruction samples using Qwen3-TTS and Seed-VC under strict quality filtration. Second, we propose a Context-Aware Direct Preference Optimization (CA-DPO) method. By employing a cascaded filtering strategy, ASR prescreening, LLM tournament ranking, and speaker similarity verification, we obtain high-confidence preference pairs that significantly enhance holistic ``Context$\rightarrow$CoT$\rightarrow$Speech'' consistency during DPO training. Third, we establish an evaluation method featuring a 500-sample test set and an LLM-as-Judge framework to independently assess reasoning and execution fidelity. Experiments demonstrate that our system significantly outperforms the baseline across all objective and subjective metrics, validating our data governance and alignment strategies.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech
Chain-of-Thought
Contextual Comprehension
Instruction Fidelity
Audio Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-Aware Direct Preference Optimization (CA-DPO)
Chain-of-Thought (CoT) reasoning
data process framework
instruction fidelity
audio quality
J
Jingbin Hu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Luyu Wang
Luyu Wang
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Wenjie Tian
Wenjie Tian
Northwest Polytechnical University
speech generation
K
Kangxiang Xia
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Q
Qirui Zhan
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, Xi’an, China
Haoyu Zhang
Haoyu Zhang
Ph.D. candidate, Norwegian University of Science and Technology
Y
Yunxiang Chen
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
H
Houdun Liu
Shenzhen Pimei Technology Co., Ltd., Guangdong, China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence
Liumeng Xue
Liumeng Xue
Hong Kong University of Science and Technology
Audio Speech and Language ProcessingSpeech Generation