MUSCAT: MUltilingual, SCientific ConversATion Benchmark

πŸ“… 2026-04-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of effective evaluation benchmarks for automatic speech recognition (ASR) systems in scenarios involving multilingual code-switching, dense scientific terminology, and conversational complexity. To bridge this gap, the authors introduce the first ASR evaluation dataset derived from authentic multilingual scientific discussions, comprising recordings of multiple speakers conversing bilingually about research papers, accompanied by audio segmentation, speaker diarization, and multilingual transcripts. The work proposes a comprehensive evaluation framework that extends beyond conventional word error rate metrics to enable consistent cross-lingual performance comparison. Experimental results demonstrate a significant performance drop among state-of-the-art ASR systems on this benchmark, underscoring both its challenge and practical relevance for advancing multilingual ASR in specialized domains.

Technology Category

Application Category

πŸ“ Abstract
The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were a multilingual speaker. To create this experience, speech technology needs to address several challenges: Handling mixed multilingual input, specific vocabulary, and code-switching. However, there is currently no dataset benchmarking this situation. We propose a new benchmark to evaluate current Automatic Speech Recognition (ASR) systems, whether they are able to handle these challenges. The benchmark consists of bilingual discussions on scientific papers between multiple speakers, each conversing in a different language. We provide a standard evaluation framework, beyond Word Error Rate (WER) enabling consistent comparison of ASR performance across languages. Experimental results demonstrate that the proposed dataset is still an open challenge for state-of-the-art ASR systems. The dataset is available in https://huggingface.co/datasets/goodpiku/muscat-eval \\ \newline \Keywords{multilingual, speech recognition, audio segmentation, speaker diarization}
Problem

Research questions and friction points this paper is trying to address.

multilingual
speech recognition
code-switching
benchmark
scientific conversation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multilingual
speech recognition
code-switching
scientific conversation
evaluation benchmark
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
S
Supriti Sinhamahapatra
Karlsruhe Institute of Technology
T
Thai-Binh Nguyen
Karlsruhe Institute of Technology
Y
Yiğit Oğuz
Karlsruhe Institute of Technology
E
Enes Ugan
Karlsruhe Institute of Technology
Jan Niehues
Jan Niehues
Institute for Anthropomatics and Robotics (IAR), Karlsruhe Institute for Technology (KIT)
natural language processing - machine translation
Alexander Waibel
Alexander Waibel
Carnegie Mellon, KIT, Karlsruhe Institute of Technology, University of Karlsruhe
Machine LearningNeural NetworksSpeech TranslationMultimodal Interfaces