Can we train ASR systems on Code-switch without real code-switch data? Case study for Singapore's languages
This study addresses the challenge of code-switching (CS) automatic speech recognition (ASR) for low-resource language pairs in Singapore—Malay–English, Mandarin–Malay, and Tamil–English—using monolingual data only. We propose a novel phrase-level natural-pattern CS data synthesis method that requires no authentic CS annotations. Leveraging this approach, we construct the first comprehensive CS-ASR benchmark covering Southeast Asian multilingual scenarios. By fine-tuning large pre-trained models—including Whisper, MMS, and SeamlessM4T—with our synthetic CS data and monolingual data augmentation, we achieve significant improvements in both CS and monolingual ASR performance, with the largest gains observed for Malay–English. Our work establishes a cost-effective, high-fidelity, purely synthetic-data-driven paradigm for low-resource CS-ASR, eliminating reliance on scarce and expensive annotated CS speech data.