SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决东南亚语言在语音理解评估中的不足,通过SEA-SpeechBench基准,使用97,194个样本和597小时音频数据进行多任务评估。
📝 Abstract
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leaving Southeast Asian (SEA) languages critically underrepresented. We introduce SEA-SpeechBench, to the best of our knowledge, the first large-scale multitask benchmark that evaluates speech understanding in 11 SEA languages through 97,194 samples across 99 evaluation sets and 597 hours of curated audio data. Our benchmark comprises 9 diverse tasks across 3 categories: speech processing (automatic speech recognition, speech translation, spoken question answering), paralinguistic analysis (emotion, gender, age, speaker recognition), and temporal understanding, a novel dimension featuring timestamped content queries and temporal localization within extended audio sequences up to 3 minutes. We implement multilingual prompting in both native SEA languages and English to reflect user interactions with audio-language models. Evaluation of leading open-source and proprietary systems reveals marked performance gaps. Across all models, performance remains underwhelming on temporal understanding, emotion recognition, and speech translation. Prompting in low-resource languages such as Burmese and Tamil lags behind English by up to 41 percentage points. Our findings expose critical model limitations and underscore the need for inclusive model development. The SEA-SpeechBench benchmark is available at https://zwenyu.github.io/SEA-SpeechBench/.
Problem

Research questions and friction points this paper is trying to address.

Speech Understanding
Southeast Asian Languages
Evaluation Frameworks
Multitask Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

multitask benchmark
speech understanding
Southeast Asian languages
temporal understanding
multilingual prompting
🔎 Similar Papers
2024-06-14Conference on Empirical Methods in Natural Language ProcessingCitations: 6
💼 Related Jobs
No related jobs found.
J
Jingyi Liao
Institute of Advanced Intelligence and Computing, A*STAR; Nanyang Technological University
Wenyu Zhang
Wenyu Zhang
AI Researcher, Center for AI Safety
Machine LearningAIMultimodal AI
Zhuohan Liu
Zhuohan Liu
Research Engineer
Y
Yingxu He
Institute of Advanced Intelligence and Computing, A*STAR
Geyu Lin
Geyu Lin
Research Engineer, I2R, A*STAR
Generative AINLPSpeech
X
Xunlong Zou
Institute of Advanced Intelligence and Computing, A*STAR
Shuo Sun
Shuo Sun
Johns Hopkins University
S
Syed Ali Redha Alsagoff
Nanyang Technological University
A
Ai Ti Aw
Institute of Advanced Intelligence and Computing, A*STAR