Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
This work addresses the challenge of developing an efficient and low-cost automatic speech recognition (ASR) system for Singapore’s multilingual setting—encompassing English, Mandarin, Malay, and Tamil—by proposing a balanced sampling fine-tuning strategy that operates without explicit language labels. The approach enables end-to-end training of Qwen3-ASR-0.6B/1.7B models to perform implicit language identification and transcription jointly. The resulting compact multilingual ASR system, Polyglot-Lion-1.7B, achieves an average word error rate of 14.85% across 12 benchmarks, with a remarkably low training cost of just \$81 on a single GPU and an inference speed of 0.10 seconds per sample—approximately 20 times faster than MERaLiON. Despite its significantly reduced model size and computational overhead, the system approaches the performance of large, specialized ASR systems.