Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了无监督语音模型对不同口音适应的问题,通过引入ABX-Accent基准和使用自适应域归一化方法微调预训练模型来改善跨说话者ABX得分。
📝 Abstract
Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. Yet, little is known about how such methods deal with out-of-domain speech and how could they be adapted in a few shot to new domains. This is important especially for accented speech where one observes a long tail of accents that diverge from the standard ones. We introduce ABX- Accent, a benchmark based on the AESRC dataset that features 10 different accents of English. It includes a small (< 10 hours) unlabelled training set in each of the accents and adaptations of the Zero Resources Challenge ABX evaluation metrics to each of the accents. We illustrate this benchmark with a baseline model that uses adaptive domain normalization to fine tune a pretrained Contrastive Predictive Coding model on the accents. This method is first developed on LibriSpeech using a male/female split. When applied to the new benchmark, the proposed method yields a relative improvement of 23.6% on across-speaker ABX scores on average compared to non adapted models. The data and metrics will be open sourced upon paper acceptance
Problem

Research questions and friction points this paper is trying to address.

unsupervised speech modeling
domain adaptation
accented speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

ABX-Accent
Adaptive Domain Normalization
Contrastive Predictive Coding
🔎 Similar Papers