A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究开发了低资源语言Mizo的自动语音识别系统,通过收集17.62小时语音数据并使用Whisper和SraVaani 1.0模型进行微调,显著降低了词错误率。
📝 Abstract
This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.
Problem

Research questions and friction points this paper is trying to address.

Automatic Speech Recognition
low-resource language
Mizo
speech data
fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatic Speech Recognition
Mizo
Whisper model
morphology-aware evaluation
low-resource language