FLEURS-Kobani: Extending the FLEURS Dataset for Northern Kurdish

📅 2026-03-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of publicly available speech benchmarks for Northern Kurdish (Kurmanji, KMR), which has hindered progress in automatic speech recognition (ASR) and speech translation. To bridge this gap, the authors introduce FLEURS-Kobani, the first KMR-focused dataset derived from the FLEURS benchmark, comprising 5,162 validated utterances (18 hours and 24 minutes) from 31 native speakers, enabling research in ASR, end-to-end speech-to-text translation (S2TT), and speech-to-speech translation (S2ST). Leveraging the Whisper v3-large architecture with Common Voice pretraining and a two-stage fine-tuning strategy, the proposed system achieves a word error rate (WER) of 28.11% and character error rate (CER) of 9.84% on ASR, and a BLEU score of 8.68 for end-to-end KMR→EN S2TT, establishing a foundational benchmark for low-resource Kurdish speech processing.

Technology Category

Application Category

📝 Abstract
FLEURS offers n-way parallel speech for 100+ languages, but Northern Kurdish is not one of them, which limits benchmarking for automatic speech recognition and speech translation tasks in this language. We present FLEURS-Kobani, a Northern Kurdish (ISO 639-3 KMR) spoken extension of the FLEURS benchmark. The FLEURS-Kobani dataset consists of 5,162 validated utterances, totaling 18 hours and 24 minutes. The data were recorded by 31 native speakers. It extends benchmark coverage to an under-resourced Kurdish variety. As baselines, we fine-tuned Whisper v3-large for ASR and E2E S2TT. A two-stage fine-tuning strategy (Common Voice to FLEURS-Kobani) yields the best ASR performance (WER 28.11, CER 9.84 on test). For E2E S2TT (KMR to EN), Whisper achieves 8.68 BLEU on test; we additionally report pivot-derived targets and a cascaded S2TT setup. FLEURS-Kobani provides the first public Northern Kurdish benchmark for evaluation of ASR, S2TT and S2ST tasks. The dataset is publicly released for research use under a CC BY 4.0 license.
Problem

Research questions and friction points this paper is trying to address.

Northern Kurdish
automatic speech recognition
speech translation
under-resourced language
benchmark dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

FLEURS-Kobani
Northern Kurdish
low-resource speech
two-stage fine-tuning
speech translation benchmark
🔎 Similar Papers
D
Daban Q. Jaff
1Erfurt University, Erfurt, Germany; 2Koya University, Koysinjaq, Iraq
M
Mohammad Mohammadamini
3LIUM, Le Mans University, Le Mans, France