ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks
This work addresses the NADI multilingual Arabic dialect speech processing task, tackling two core challenges: Arabic Dialect Identification (ADI) and multilingual Automatic Speech Recognition (ASR). We propose a joint optimization framework based on large-model fine-tuning and dialect-specific data augmentation. For ADI, we employ the Whisper-large-v3 encoder with dialect-aware data augmentation to achieve end-to-end dialect classification. For ASR, we fine-tune the SeamlessM4T-v2 Large model separately on each of eight Arabic dialects to enhance cross-dialect robustness. Our approach significantly outperforms baselines: achieving 79.83% accuracy on ADI (ranked first), and average WER/CER of 38.54%/14.53% on ASR (ranked second). The key contribution lies in empirically validating that dialect-specific fine-tuning combined with domain-adaptive data augmentation substantially improves low-resource multilingual speech modeling performance.