FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish
This study addresses the scarcity of publicly available speech benchmarks for Northern Kurdish (Kurmanji, KMR), which has hindered progress in automatic speech recognition (ASR) and speech translation. To bridge this gap, the authors introduce FLEURS-Kobani, the first KMR-focused dataset derived from the FLEURS benchmark, comprising 5,162 validated utterances (18 hours and 24 minutes) from 31 native speakers, enabling research in ASR, end-to-end speech-to-text translation (S2TT), and speech-to-speech translation (S2ST). Leveraging the Whisper v3-large architecture with Common Voice pretraining and a two-stage fine-tuning strategy, the proposed system achieves a word error rate (WER) of 28.11% and character error rate (CER) of 9.84% on ASR, and a BLEU score of 8.68 for end-to-end KMR→EN S2TT, establishing a foundational benchmark for low-resource Kurdish speech processing.