Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
Low-resource languages like Thai face critical challenges in speech large language modeling (SLLM), including poor speech encoder performance, weak multimodal understanding capabilities, high computational cost of ASR-based forced alignment, and scarcity of paired speech-text data. To address these, this work proposes a systematic solution: (1) the first Thai self-supervised speech encoder, XLSR-Thai; (2) U-Align, a lightweight cross-modal alignment method that replaces expensive ASR-based forced alignment; and (3) Thai-SUP, a scalable Thai understanding data synthesis framework generating over 1,000 hours of high-quality, multitask training data. Through joint optimization via self-supervised pretraining, U-Align fine-tuning, and cross-lingual transfer, our approach significantly improves Thai speech recognition, semantic understanding, and instruction-following performance. All models and datasets are publicly released, establishing essential infrastructure for low-resource speech understanding research.