LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

📅 2025-04-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Tunisian Arabic speech recognition faces a low-resource bottleneck due to scarce annotated speech data and inadequate modeling of English–French code-switching. To address this, we introduce LinTO—the first high-quality, open-source ASR dataset for Tunisian Arabic—comprising multi-source text, real-world speech recordings (including natural code-switching), and phoneme-level aligned transcriptions. We propose a dialect-specific phonological annotation schema and an explicit code-switching modeling framework. Leveraging systematic speech collection, multilingual alignment, audio augmentation, and a phonology-informed evaluation methodology, LinTO contains tens of thousands of precisely annotated utterances, filling a critical benchmark resource gap. Evaluated on standard test sets, ASR models trained on LinTO achieve a 22% relative reduction in word error rate (WER). LinTO thus establishes the first authoritative ASR benchmark and end-to-end technical solution for Tunisian Arabic.

Technology Category

Application Category

📝 Abstract
Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO audio and textual datasets -- comprehensive resources that capture phonological and lexical features of Tunisian Arabic Dialect. These datasets include a variety of texts from numerous sources and real-world audio samples featuring diverse speakers and code-switching between Tunisian Arabic Dialect and English or French. By providing high-quality audio paired with precise transcriptions, the LinTO audio and textual datasets aim to provide qualitative material to build and benchmark ASR systems for the Tunisian Arabic Dialect. Keywords -- Tunisian Arabic Dialect, Speech-to-Text, Low-Resource Languages, Audio Data Augmentation
Problem

Research questions and friction points this paper is trying to address.

Addressing scarcity of annotated speech datasets for Tunisian Arabic Dialect
Capturing phonological and lexical features in diverse speech samples
Providing high-quality audio-transcription pairs for ASR benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive datasets for Tunisian Arabic Dialect
Diverse audio samples with code-switching
High-quality audio paired with transcriptions
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hedi Naouara
LINAGORA
J
Jean-Pierre Lorr'e
LINAGORA
J
Jérôme Louradour
LINAGORA