IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

πŸ“… 2026-06-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing evaluation methods struggle to discern whether audio large language models (AudioLLMs) genuinely leverage input context or merely rely on pre-trained knowledge. To address this gap, this work introduces a multilingual speech benchmark comprising 56 hours of audio across eight Indian languages and 23 specialized domains, alongside a novel seven-level contextual prompting framework. This framework incrementally incorporates signals such as metadata, natural language descriptions, entity lists, and adversarial prompts to systematically assess models’ contextual grounding capabilities. Experiments on five prominent AudioLLMs reveal substantial differences in their ability to utilize provided context, underscoring the necessity of explicit evaluation protocols and filling a critical void in the current assessment landscape for audio-language models.
πŸ“ Abstract
AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.
Problem

Research questions and friction points this paper is trying to address.

Audio Large Language Models
context utilisation
speech recognition
benchmarking
contextual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

context utilisation
Audio Large Language Models
multilingual benchmark
prompting framework
contextual grounding
πŸ’Ό Related Jobs
No related jobs found.
S
Sakshi Joshi
AI4Bharat, Indian Institute of Technology Madras, India
D
Dhruv Subhash Rathi
Sarvam AI, India
S
Sanskar Singh
Sarvam AI, India
E
Eldho Ittan George
AI4Bharat, Indian Institute of Technology Madras, India
R
R J Hari
AI4Bharat, Indian Institute of Technology Madras, India
K
Kaushal Bhogale
AI4Bharat, Indian Institute of Technology Madras, India
M
Mitesh M. Khapra
AI4Bharat, Indian Institute of Technology Madras, India