Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建统一的基于音素的TTS增强管道及提出音素频率引导选择方法,解决自动语音识别中合成语音数据的有效利用问题。
📝 Abstract
Synthetic speech provides scalable supervision for automatic speech recognition (ASR), but its benefit depends on the selected texts, reference speech, and amount of synthesized data. We present a unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning. The pipeline combines language-specific grapheme-to-phoneme conversion, reference-speech filtering, candidate-text selection, synthesis, and matched ASR continuation. We further propose phoneme-frequency-guided selection (PFGS), which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels. Experiments with separate monolingual ASR systems for Arabic, French, Italian, and Portuguese span 13 test sets. Across the synthesis-scale sweep, random augmentation improves over matched real-only continuation on 11 test sets. Under a nominal 60% synthesis budget, PFGS improves over real-only training on 12 test sets and over random selection on 9. Its largest relative word error rate (WER) reduction against random selection is 19.3%. With target texts and synthesis counts fixed, reference-speech filtering reduces absolute WER by 0.29 and 0.59 points on Italian and French Common Voice, respectively. These results identify synthesis scale, candidate-text content, and reference quality as important control variables in TTS-based ASR augmentation.
Problem

Research questions and friction points this paper is trying to address.

TTS augmentation
ASR
synthetic speech
candidate texts
reference speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

phoneme-based TTS-to-ASR augmentation
multilingual TTS model
F5-TTS architecture
phoneme-frequency-guided selection (PFGS)
🔎 Similar Papers
No similar papers found.
Z
Zhen Wang
Shanghai Qi Zhi Institute, Shanghai, China
T
TianRui Wu
Shanghai Qi Zhi Institute, Shanghai, China
R
RongQi Han
Shanghai Qi Zhi Institute, Shanghai, China
H
Hao Wu
Shanghai Qi Zhi Institute, Shanghai, China
W
Wei Liang
Megatronix (Beijing) Technology Co., Ltd.