ADI-20: Arabic Dialect Identification dataset and models
This work addresses Arabic Dialect Identification (ADI), a challenging multilingual speech classification task. We introduce ADI-20, the first large-scale, publicly available dataset covering all 22 Arab League countries’ dialects plus Modern Standard Arabic (MSA), comprising 19 dialects and 3,556 hours of speech. We propose an end-to-end ADI model built upon the ECAPA-TDNN backbone, enhanced with Whisper encoder blocks, attention-based pooling, and a dialect-specific classification head. To our knowledge, this is the first open, reproducible framework for pan-Arabic dialect modeling—releasing data, models, and training code. Experiments demonstrate state-of-the-art performance, exceptional data efficiency (F1 drops by <1.5% when trained on only 30% of the data), and systematic empirical analysis of the scaling relationships between dataset size, model parameters, and ADI accuracy.