Institution profile

Emotech Labs

Industry researcheurope · gb
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CTC-DID: CTC-Based Arabic dialect identification for streaming applications

Jan 18, 2026

This work addresses the challenge of insufficient accuracy and real-time performance in Arabic dialect identification under low-resource, streaming conditions. The authors propose a limited-vocabulary speech recognition framework based on Connectionist Temporal Classification (CTC) loss, which models dialect labels as sequences of phonetic units and enables end-to-end streaming inference. This study is the first to apply CTC to dialect identification and introduces a language-agnostic heuristic label repetition strategy that significantly enhances robustness for short utterances and zero-shot scenarios. By integrating self-supervised learning (SSL) models with CTC loss and leveraging alignment labels generated via LAH or pretrained ASR systems, the proposed approach outperforms fine-tuned Whisper and ECAPA-TDNN baselines on low-resource Arabic dialect recognition tasks, achieving superior performance particularly on the Casablanca dataset in zero-shot and short-duration evaluations.

0 citationsRead paper

SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition

Jun 27, 2025

To address performance degradation of foundation models in low-resource Arabic–English code-switching (CS) speech recognition due to data scarcity, this paper proposes SAGE—a speech synthesis method for audio splicing—and an experience-replay-inspired incremental fine-tuning strategy. SAGE generates high-fidelity artificial CS speech via controllable dialect alignment and cross-lingual prosody modeling. The experience replay mechanism mitigates catastrophic forgetting during few-shot adaptation. Integrated with a self-supervised speech model, 3-gram language model fusion, and few-shot adaptation, the approach achieves a word error rate (WER) of 31.1% on the Arabic–English CS benchmark—reducing absolute WER by 5.5% and 8.4% over USM and Whisper-large-v2, respectively. The method significantly enhances robustness to Arabic dialectal variation and generalization to CS phenomena.

0 citationsRead paper
Recent publications

Latest Papers

CTC-DID: CTC-Based Arabic dialect identification for streaming applications

Jan 18, 2026

This work addresses the challenge of insufficient accuracy and real-time performance in Arabic dialect identification under low-resource, streaming conditions. The authors propose a limited-vocabulary speech recognition framework based on Connectionist Temporal Classification (CTC) loss, which models dialect labels as sequences of phonetic units and enables end-to-end streaming inference. This study is the first to apply CTC to dialect identification and introduces a language-agnostic heuristic label repetition strategy that significantly enhances robustness for short utterances and zero-shot scenarios. By integrating self-supervised learning (SSL) models with CTC loss and leveraging alignment labels generated via LAH or pretrained ASR systems, the proposed approach outperforms fine-tuned Whisper and ECAPA-TDNN baselines on low-resource Arabic dialect recognition tasks, achieving superior performance particularly on the Casablanca dataset in zero-shot and short-duration evaluations.

0 citationsRead paper

SAGE: Spliced-Audio Generated Data for Enhancing Foundational Models in Low-Resource Arabic-English Code-Switched Speech Recognition

Jun 27, 2025

To address performance degradation of foundation models in low-resource Arabic–English code-switching (CS) speech recognition due to data scarcity, this paper proposes SAGE—a speech synthesis method for audio splicing—and an experience-replay-inspired incremental fine-tuning strategy. SAGE generates high-fidelity artificial CS speech via controllable dialect alignment and cross-lingual prosody modeling. The experience replay mechanism mitigates catastrophic forgetting during few-shot adaptation. Integrated with a self-supervised speech model, 3-gram language model fusion, and few-shot adaptation, the approach achieves a word error rate (WER) of 31.1% on the Arabic–English CS benchmark—reducing absolute WER by 5.5% and 8.4% over USM and Whisper-large-v2, respectively. The method significantly enhances robustness to Arabic dialectal variation and generalization to CS phenomena.

0 citationsRead paper