π€ AI Summary
Existing evaluation methods struggle to discern whether audio large language models (AudioLLMs) genuinely leverage input context or merely rely on pre-trained knowledge. To address this gap, this work introduces a multilingual speech benchmark comprising 56 hours of audio across eight Indian languages and 23 specialized domains, alongside a novel seven-level contextual prompting framework. This framework incrementally incorporates signals such as metadata, natural language descriptions, entity lists, and adversarial prompts to systematically assess modelsβ contextual grounding capabilities. Experiments on five prominent AudioLLMs reveal substantial differences in their ability to utilize provided context, underscoring the necessity of explicit evaluation protocols and filling a critical void in the current assessment landscape for audio-language models.
π Abstract
AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.