Stuttering-Aware Automatic Speech Recognition for Indonesian Language

📅 2026-01-07
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation of automatic speech recognition (ASR) systems when processing stuttered speech, a challenge particularly acute in low-resource languages like Indonesian due to the scarcity of authentic stuttered speech data. To tackle this issue, the work proposes the first stutter-aware ASR system for Indonesian, introducing a synthetic data augmentation framework that does not require real stuttered recordings. The approach generates stuttered text—featuring repetitions and prolongations—through rule-based transformations and large language models, then synthesizes corresponding stuttered speech using text-to-speech systems. This synthetic data is used to fine-tune a pretrained Whisper model. Experimental results demonstrate that the method substantially reduces word error rates on stuttered speech while preserving recognition accuracy on fluent speech, thereby enhancing the inclusivity of ASR systems for low-resource languages.

Technology Category

Application Category

📝 Abstract
Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource languages like Indonesian where specialized datasets are virtually non-existent. To overcome this scarcity, we propose a data augmentation framework that generates synthetic stuttered audio by injecting repetitions and prolongations into fluent text through a combination of rule-based transformations and large language models followed by text-to-speech synthesis. We apply this synthetic data to fine-tune a pre-trained Indonesian Whisper model using transfer learning, enabling the architecture to adapt to dysfluent acoustic patterns without requiring large-scale real-world recordings. Our experiments demonstrate that this targeted synthetic exposure consistently reduces recognition errors on stuttered speech while maintaining performance on fluent segments, validating the utility of synthetic data pipelines for developing more inclusive speech technologies in under-represented languages.
Problem

Research questions and friction points this paper is trying to address.

stuttering
automatic speech recognition
low-resource languages
Indonesian language
dysfluent speech
Innovation

Methods, ideas, or system contributions that make the work stand out.

stuttering-aware ASR
synthetic data augmentation
low-resource language
transfer learning
text-to-speech synthesis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fadhil Muhammad
Faculty of Computer Science, Universitas Indonesia
A
Alwin Djuliansah
Faculty of Computer Science, Universitas Indonesia
A
Adrian Aryaputra Hamzah
Faculty of Computer Science, Universitas Indonesia
K
Kurniawati Azizah
Faculty of Computer Science, Universitas Indonesia