🤖 AI Summary
This study addresses the challenges of speech synthesis and digital preservation for Efik, a low-resource African tonal language. The authors present the first end-to-end text-to-speech (TTS) system for Efik, leveraging a newly collected 3-hour single-speaker speech corpus. They establish reproducible TTS baselines using VITS, MMS-TTS, SpeechT5, and Orpheus-TTS, and evaluate performance through subjective assessments including MOS, Nat-MOS, and A-MOS. Among the models, MMS-TTS achieves the highest quality (MOS: 3.80 ± 0.63) and demonstrates greater stability in synthesizing long utterances, though it still exhibits tonal inaccuracies. This work provides the first systematic evaluation framework for TTS in low-resource tonal languages and underscores the need for larger-scale corpora and improved tonal modeling to advance synthesis quality.
📝 Abstract
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.