LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
This work addresses the limitations of current generative speech models in zero-shot multilingual synthesis and editing, which stem from the scarcity of large-scale, high-quality multilingual speech data with word-level timestamps. To overcome this, the authors introduce LEMAS-Dataset, an open-source corpus spanning 10 languages and 150,000 hours of speech, uniquely annotated with word-level alignment timestamps. Leveraging this dataset, they propose LEMAS-TTS, a non-autoregressive model for zero-shot multilingual text-to-speech synthesis, and LEMAS-Edit, an autoregressive model that formulates speech editing as a masked token infilling task. Through accent adversarial training, CTC loss, and adaptive decoding strategies, the models achieve substantial improvements in cross-lingual accent robustness and naturalness at edit boundaries, demonstrating the efficacy of both the dataset and the proposed methodologies.