Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages

📅 2025-06-17
📈 Citations: 1
Influential: 1
📄 PDF
🤖 AI Summary
This study addresses the challenge of code-switching (CS) automatic speech recognition (ASR) for low-resource language pairs in Singapore—Malay–English, Mandarin–Malay, and Tamil–English—using monolingual data only. We propose a novel phrase-level natural-pattern CS data synthesis method that requires no authentic CS annotations. Leveraging this approach, we construct the first comprehensive CS-ASR benchmark covering Southeast Asian multilingual scenarios. By fine-tuning large pre-trained models—including Whisper, MMS, and SeamlessM4T—with our synthetic CS data and monolingual data augmentation, we achieve significant improvements in both CS and monolingual ASR performance, with the largest gains observed for Malay–English. Our work establishes a cost-effective, high-fidelity, purely synthetic-data-driven paradigm for low-resource CS-ASR, eliminating reliance on scarce and expensive annotated CS speech data.

Technology Category

Application Category

📝 Abstract
Code-switching (CS), common in multilingual settings, presents challenges for ASR due to scarce and costly transcribed data caused by linguistic complexity. This study investigates building CS-ASR using synthetic CS data. We propose a phrase-level mixing method to generate synthetic CS data that mimics natural patterns. Utilizing monolingual augmented with synthetic phrase-mixed CS data to fine-tune large pretrained ASR models (Whisper, MMS, SeamlessM4T). This paper focuses on three under-resourced Southeast Asian language pairs: Malay-English (BM-EN), Mandarin-Malay (ZH-BM), and Tamil-English (TA-EN), establishing a new comprehensive benchmark for CS-ASR to evaluate the performance of leading ASR models. Experimental results show that the proposed training strategy enhances ASR performance on monolingual and CS tests, with BM-EN showing highest gains, then TA-EN and ZH-BM. This finding offers a cost-effective approach for CS-ASR development, benefiting research and industry.
Problem

Research questions and friction points this paper is trying to address.

Train ASR systems on code-switching without real data
Generate synthetic code-switch data mimicking natural patterns
Evaluate ASR performance on under-resourced Southeast Asian languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic phrase-level mixing for CS data
Fine-tuning pretrained ASR with synthetic data
New benchmark for under-resourced language pairs
💼 Related Jobs
No related jobs found.
T
Tuan Nguyen
Institute for Infocomm Research (I2R), A*STAR, Singapore
H
Huy-Dat Tran
Institute for Infocomm Research (I2R), A*STAR, Singapore