DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenges of data scarcity and misalignment between semantic representations and fixed acoustic supervision in end-to-end spoken dialogue modeling for low-resource Chinese dialects. To overcome these issues, the authors construct a scalable data pipeline for dialectal spoken dialogue synthesis and propose a two-stage post-training strategy incorporating a self-aligned speech supervision mechanism that dynamically aligns acoustic targets with the model’s evolving semantic representations. This approach achieves the first successful end-to-end spoken dialogue modeling for low-resource Chinese dialects, significantly outperforming existing baselines across multiple dialects. Substantial improvements are observed in dialect consistency, response quality, and speech intelligibility. The complete framework is open-sourced to facilitate future research in this underexplored domain.
πŸ“ Abstract
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.
Problem

Research questions and friction points this paper is trying to address.

low-resource dialects
end-to-end speech dialogue
semantic inconsistency
speech stability
Chinese dialects
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end speech dialogue
low-resource dialects
self-aligned speech supervision
semantic consistency
scalable data pipeline
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yi Shu
School of Artificial Intelligence, University of Chinese Academy of Sciences
T
Tianyu Peng
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Wuhan AI Research
Y
Yingzhuo Deng
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences
Wen Yang
Wen Yang
Professor, School of Electronic Information, Wuhan University
Image ProcessingPattern RecognitionMachine Learning
J
Jun Lin
GWM AI Lab
C
Changming Xie
GWM AI Lab
X
Xinyu Yu
GWM AI Lab
Jiajun Zhang
Jiajun Zhang
Institute of Automation Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelsMultimodal Information Processing