RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
This work addresses the issue of word skipping and repetition in flow-matching-based text-to-speech (TTS) synthesis, which arises from inaccurate alignments. The authors propose a latent-space augmentation strategy that explicitly models failure modes without requiring external aligners or preference data, while preserving the original input length. This approach is integrated into a contrastive flow-matching framework and represents the first application of augmentation-based contrastive flow matching to enhance TTS content fidelity. It seamlessly fits into existing zero-shot TTS pipelines. Experimental results demonstrate consistent improvements: on Seed-TTS-eval, the word error rate (WER) decreases from 1.44% to 1.38%; on the ZERO500 benchmark, character error rates (CER) for English and Korean drop from 0.48% and 0.81% to 0.35% and 0.57%, respectively, with 24 function evaluations (NFE).