Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
Existing foundation models exhibit zero-shot word error rates (WER) exceeding 100% on Southern Bantu speech recognition tasks, rendering them impractical. This work proposes a tone-conditioned curriculum learning framework that, for the first time, integrates tonal information into the curriculum mechanism. By combining mixed-difficulty scoring, tone-statistics-gated adapters, and a staged training strategy, the approach significantly enhances cross-lingual transfer performance for low-resource Bantu languages. Experiments on W2V-BERT and Whisper demonstrate that the method reduces average WER to 28.41% across six languages, achieving 23.79% on the Xitsonga transfer task. The study further uncovers systematic interactions between model architecture and language family, offering an effective adaptive training paradigm for low-resource speech recognition.