Institution profile

Elyadata

Industry research
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

Nov 13, 2025Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks

This work addresses the NADI multilingual Arabic dialect speech processing task, tackling two core challenges: Arabic Dialect Identification (ADI) and multilingual Automatic Speech Recognition (ASR). We propose a joint optimization framework based on large-model fine-tuning and dialect-specific data augmentation. For ADI, we employ the Whisper-large-v3 encoder with dialect-aware data augmentation to achieve end-to-end dialect classification. For ASR, we fine-tune the SeamlessM4T-v2 Large model separately on each of eight Arabic dialects to enhance cross-dialect robustness. Our approach significantly outperforms baselines: achieving 79.83% accuracy on ADI (ranked first), and average WER/CER of 38.54%/14.53% on ASR (ranked second). The key contribution lies in empirically validating that dialect-specific fine-tuning combined with domain-adaptive data augmentation substantially improves low-resource multilingual speech modeling performance.

1 citations1 influentialRead paper

CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech

Apr 02, 2026

This study addresses the challenge of end-to-end named entity recognition (NER) from Arabic speech, which is hindered by morphological complexity, vowel omission, and scarce annotated resources. The authors introduce CV-18 NER, the first publicly available Arabic speech NER dataset, annotated with the fine-grained Wojood schema comprising 21 entity types. They systematically evaluate both pipeline and end-to-end approaches, demonstrating that end-to-end models significantly outperform traditional cascaded systems. Specifically, AraBEST-RQ (300M), leveraging Arabic self-supervised pretraining, achieves a CoER of 37.0%, while Whisper-medium attains 38.0% CVER. The work further highlights the efficacy of multilingual weak supervision for cross-lingual transfer in low-resource settings. The dataset and code are publicly released to support future research.

0 citationsRead paper

Ara-Best-RQ: Multi Dialectal Arabic SSL

Mar 23, 2026

This work addresses the long-standing scarcity of efficient, dialect-specific self-supervised models for multilingual Arabic speech processing. It presents the first self-supervised pretraining approach tailored to the diverse family of Arabic dialects, leveraging the Conformer architecture within the BEST-RQ framework and pretrained on 5,640 hours of web-crawled and publicly available speech data. The resulting model achieves state-of-the-art performance in dialect identification with fewer parameters and significantly outperforms existing general-purpose multilingual or non-Arabic monolingual models on automatic speech recognition tasks. These results demonstrate the effectiveness and superiority of domain-customized pretraining for low-resource, linguistically heterogeneous language varieties such as Arabic dialects.

0 citationsRead paper

SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding

Mar 23, 2026

This work addresses the scarcity of high-quality spoken language understanding resources for Tunisian Arabic, which has severely hindered its deployment in task-oriented dialogue systems. To bridge this gap, the authors introduce SLURP-TN, the first spoken language understanding dataset for Tunisian dialect, comprising approximately five hours of speech from 55 native speakers across six domains, totaling 4,165 annotated utterances. Data quality is ensured through a combination of human translation and native-speaker recordings. The study further develops baseline automatic speech recognition (ASR) and spoken language understanding (SLU) systems leveraging deep neural networks and pretrained language models. Both the dataset and associated models are publicly released, filling a critical void in structured SLU benchmarks for low-resource dialects and significantly advancing research and applications of Tunisian Arabic in spoken interactive systems.

0 citationsRead paper

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

Nov 13, 2025Proceedings of The Third Arabic Natural Language Processing Conference

Data scarcity severely hampers Tunisian Arabic–English speech translation, particularly due to the absence of publicly available code-switched speech corpora, impeding low-resource dialectal NLP research. To address this, we introduce TEDxTN—the first open-source, code-switched Tunisian Arabic–English speech translation corpus—covering speakers from 11+ regions in Tunisia, 108 TEDx talks, and 25 hours of audio. The corpus features professional segmentation, bilingual transcription, human translation, and a publicly released annotation guideline. Methodologically, we present the first end-to-end joint modeling framework for automatic speech recognition and speech translation, leveraging pretrained models fine-tuned on TEDxTN to establish a strong, reproducible baseline. TEDxTN fills a critical gap in Arabic dialectal speech translation resources and serves as a foundational benchmark for low-resource spoken language translation, code-switching modeling, and dialectal NLP.

0 citationsRead paper
Recent publications

Latest Papers

CV-18 NER: Augmented Common Voice for Named Entity Recognition from Arabic Speech

Apr 02, 2026

This study addresses the challenge of end-to-end named entity recognition (NER) from Arabic speech, which is hindered by morphological complexity, vowel omission, and scarce annotated resources. The authors introduce CV-18 NER, the first publicly available Arabic speech NER dataset, annotated with the fine-grained Wojood schema comprising 21 entity types. They systematically evaluate both pipeline and end-to-end approaches, demonstrating that end-to-end models significantly outperform traditional cascaded systems. Specifically, AraBEST-RQ (300M), leveraging Arabic self-supervised pretraining, achieves a CoER of 37.0%, while Whisper-medium attains 38.0% CVER. The work further highlights the efficacy of multilingual weak supervision for cross-lingual transfer in low-resource settings. The dataset and code are publicly released to support future research.

0 citationsRead paper

Ara-Best-RQ: Multi Dialectal Arabic SSL

Mar 23, 2026

This work addresses the long-standing scarcity of efficient, dialect-specific self-supervised models for multilingual Arabic speech processing. It presents the first self-supervised pretraining approach tailored to the diverse family of Arabic dialects, leveraging the Conformer architecture within the BEST-RQ framework and pretrained on 5,640 hours of web-crawled and publicly available speech data. The resulting model achieves state-of-the-art performance in dialect identification with fewer parameters and significantly outperforms existing general-purpose multilingual or non-Arabic monolingual models on automatic speech recognition tasks. These results demonstrate the effectiveness and superiority of domain-customized pretraining for low-resource, linguistically heterogeneous language varieties such as Arabic dialects.

0 citationsRead paper

SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding

Mar 23, 2026

This work addresses the scarcity of high-quality spoken language understanding resources for Tunisian Arabic, which has severely hindered its deployment in task-oriented dialogue systems. To bridge this gap, the authors introduce SLURP-TN, the first spoken language understanding dataset for Tunisian dialect, comprising approximately five hours of speech from 55 native speakers across six domains, totaling 4,165 annotated utterances. Data quality is ensured through a combination of human translation and native-speaker recordings. The study further develops baseline automatic speech recognition (ASR) and spoken language understanding (SLU) systems leveraging deep neural networks and pretrained language models. Both the dataset and associated models are publicly released, filling a critical void in structured SLU benchmarks for low-resource dialects and significantly advancing research and applications of Tunisian Arabic in spoken interactive systems.

0 citationsRead paper

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

Nov 13, 2025Proceedings of The Third Arabic Natural Language Processing Conference

Data scarcity severely hampers Tunisian Arabic–English speech translation, particularly due to the absence of publicly available code-switched speech corpora, impeding low-resource dialectal NLP research. To address this, we introduce TEDxTN—the first open-source, code-switched Tunisian Arabic–English speech translation corpus—covering speakers from 11+ regions in Tunisia, 108 TEDx talks, and 25 hours of audio. The corpus features professional segmentation, bilingual transcription, human translation, and a publicly released annotation guideline. Methodologically, we present the first end-to-end joint modeling framework for automatic speech recognition and speech translation, leveraging pretrained models fine-tuned on TEDxTN to establish a strong, reproducible baseline. TEDxTN fills a critical gap in Arabic dialectal speech translation resources and serves as a foundational benchmark for low-resource spoken language translation, code-switching modeling, and dialectal NLP.

0 citationsRead paper

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

Nov 13, 2025Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks

This work addresses the NADI multilingual Arabic dialect speech processing task, tackling two core challenges: Arabic Dialect Identification (ADI) and multilingual Automatic Speech Recognition (ASR). We propose a joint optimization framework based on large-model fine-tuning and dialect-specific data augmentation. For ADI, we employ the Whisper-large-v3 encoder with dialect-aware data augmentation to achieve end-to-end dialect classification. For ASR, we fine-tune the SeamlessM4T-v2 Large model separately on each of eight Arabic dialects to enhance cross-dialect robustness. Our approach significantly outperforms baselines: achieving 79.83% accuracy on ADI (ranked first), and average WER/CER of 38.54%/14.53% on ASR (ranked second). The key contribution lies in empirically validating that dialect-specific fine-tuning combined with domain-adaptive data augmentation substantially improves low-resource multilingual speech modeling performance.

1 citations1 influentialRead paper