Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
This work addresses the absence of speech benchmarks that comprehensively cover all 22 official languages of India and reflect real-world multilingual scenarios for joint speaker diarization and automatic speech recognition (ASR). To bridge this gap, we introduce and publicly release Indic DiarBench, a benchmark dataset comprising 108 hours of naturally occurring multi-speaker audio spanning near-field meetings, far-field recordings, and in-the-wild settings. It is the first dataset to fully encompass all Indian official languages and incorporates complex linguistic phenomena such as code-switching, dialectal variation, and speaker overlap. The dataset includes human-verified, time-aligned transcripts with speaker labels. We further establish baseline systems leveraging commercial ASR APIs and multimodal large language models, providing a standardized evaluation platform to advance research in multilingual joint diarization and ASR and foster more inclusive speech technologies.