Institution profile

AI4Bharat

Academic institutionasia · in
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

Jul 26, 2026

This work addresses the absence of speech benchmarks that comprehensively cover all 22 official languages of India and reflect real-world multilingual scenarios for joint speaker diarization and automatic speech recognition (ASR). To bridge this gap, we introduce and publicly release Indic DiarBench, a benchmark dataset comprising 108 hours of naturally occurring multi-speaker audio spanning near-field meetings, far-field recordings, and in-the-wild settings. It is the first dataset to fully encompass all Indian official languages and incorporates complex linguistic phenomena such as code-switching, dialectal variation, and speaker overlap. The dataset includes human-verified, time-aligned transcripts with speaker labels. We further establish baseline systems leveraging commercial ASR APIs and multimodal large language models, providing a standardized evaluation platform to advance research in multilingual joint diarization and ASR and foster more inclusive speech technologies.

0 citationsRead paper

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

Jun 17, 2026

Existing evaluation methods struggle to discern whether audio large language models (AudioLLMs) genuinely leverage input context or merely rely on pre-trained knowledge. To address this gap, this work introduces a multilingual speech benchmark comprising 56 hours of audio across eight Indian languages and 23 specialized domains, alongside a novel seven-level contextual prompting framework. This framework incrementally incorporates signals such as metadata, natural language descriptions, entity lists, and adversarial prompts to systematically assess models’ contextual grounding capabilities. Experiments on five prominent AudioLLMs reveal substantial differences in their ability to utilize provided context, underscoring the necessity of explicit evaluation protocols and filling a critical void in the current assessment landscape for audio-language models.

0 citationsRead paper

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

Apr 23, 2026

This work addresses the high variance and multidimensional perceptual complexity in evaluating multilingual text-to-speech (TTS) systems for Indian languages by proposing the first scalable, language-controllable evaluation framework. Through large-scale crowdsourced pairwise preference experiments across ten Indian languages—encompassing over 5,000 utterances and 120,000 judgments—the study integrates multidimensional perceptual annotations, Bradley–Terry modeling, and SHAP-based interpretability analysis. This approach reveals, for the first time, the performance trade-offs of TTS models across six key dimensions: intelligibility, expressiveness, audio quality, naturalness, speaker similarity, and linguistic accuracy. The project establishes the first human preference leaderboard for Indian multilingual TTS, validates evaluation reliability, and provides an interpretable, fine-grained benchmark for future research in multilingual speech synthesis.

0 citationsRead paper

Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models

Mar 15, 2026

This work addresses the scarcity of high-quality, domain-adapted training corpora for small language models in low-resource Indian languages. To this end, the authors propose a hybrid data construction approach that integrates local generation with cross-lingual expansion. Leveraging the Sarvam-M model within a compositional prompt engineering framework, they generate native-language content and further augment it through multilingual expansion using the Google Translate API, complemented by programmatic filtering to ensure quality. The resulting synthetic corpus spans 17 Indian languages and comprises 132,942 children’s stories—over 93.9 million tokens—in concise, narratively coherent texts strictly rendered in native scripts. This resource provides a foundational dataset for training and transfer learning of small language models targeting low-resource Indian languages.

0 citationsRead paper

Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes

Mar 15, 2026

Existing decoding strategies such as Top-k and Top-p rely on static truncation, which struggles to adapt to the dynamically varying information density in language generation, often failing to balance creativity and logical coherence. This work proposes Top-b decoding, the first approach to formalize the decoding process as an entropy-driven dynamic bandwidth modulation mechanism over a relative probability manifold. By coupling an adaptive bandwidth coefficient strictly with the model’s instantaneous Shannon entropy, Top-b dynamically adjusts the size of the candidate token set. Theoretical analysis demonstrates that Top-b minimizes tail distribution variance, thereby establishing an approximately self-regulating generative control system. Empirical results show that this method significantly reduces the variance between generation entropy and decoding behavior on benchmarks such as GPQA and GSM8K, while maintaining strong reasoning accuracy.

0 citationsRead paper
Recent publications

Latest Papers

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

Jul 26, 2026

This work addresses the absence of speech benchmarks that comprehensively cover all 22 official languages of India and reflect real-world multilingual scenarios for joint speaker diarization and automatic speech recognition (ASR). To bridge this gap, we introduce and publicly release Indic DiarBench, a benchmark dataset comprising 108 hours of naturally occurring multi-speaker audio spanning near-field meetings, far-field recordings, and in-the-wild settings. It is the first dataset to fully encompass all Indian official languages and incorporates complex linguistic phenomena such as code-switching, dialectal variation, and speaker overlap. The dataset includes human-verified, time-aligned transcripts with speaker labels. We further establish baseline systems leveraging commercial ASR APIs and multimodal large language models, providing a standardized evaluation platform to advance research in multilingual joint diarization and ASR and foster more inclusive speech technologies.

0 citationsRead paper

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

Jun 17, 2026

Existing evaluation methods struggle to discern whether audio large language models (AudioLLMs) genuinely leverage input context or merely rely on pre-trained knowledge. To address this gap, this work introduces a multilingual speech benchmark comprising 56 hours of audio across eight Indian languages and 23 specialized domains, alongside a novel seven-level contextual prompting framework. This framework incrementally incorporates signals such as metadata, natural language descriptions, entity lists, and adversarial prompts to systematically assess models’ contextual grounding capabilities. Experiments on five prominent AudioLLMs reveal substantial differences in their ability to utilize provided context, underscoring the necessity of explicit evaluation protocols and filling a critical void in the current assessment landscape for audio-language models.

0 citationsRead paper

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

Apr 23, 2026

This work addresses the high variance and multidimensional perceptual complexity in evaluating multilingual text-to-speech (TTS) systems for Indian languages by proposing the first scalable, language-controllable evaluation framework. Through large-scale crowdsourced pairwise preference experiments across ten Indian languages—encompassing over 5,000 utterances and 120,000 judgments—the study integrates multidimensional perceptual annotations, Bradley–Terry modeling, and SHAP-based interpretability analysis. This approach reveals, for the first time, the performance trade-offs of TTS models across six key dimensions: intelligibility, expressiveness, audio quality, naturalness, speaker similarity, and linguistic accuracy. The project establishes the first human preference leaderboard for Indian multilingual TTS, validates evaluation reliability, and provides an interpretable, fine-grained benchmark for future research in multilingual speech synthesis.

0 citationsRead paper

Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models

Mar 15, 2026

This work addresses the scarcity of high-quality, domain-adapted training corpora for small language models in low-resource Indian languages. To this end, the authors propose a hybrid data construction approach that integrates local generation with cross-lingual expansion. Leveraging the Sarvam-M model within a compositional prompt engineering framework, they generate native-language content and further augment it through multilingual expansion using the Google Translate API, complemented by programmatic filtering to ensure quality. The resulting synthetic corpus spans 17 Indian languages and comprises 132,942 children’s stories—over 93.9 million tokens—in concise, narratively coherent texts strictly rendered in native scripts. This resource provides a foundational dataset for training and transfer learning of small language models targeting low-resource Indian languages.

0 citationsRead paper

Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes

Mar 15, 2026

Existing decoding strategies such as Top-k and Top-p rely on static truncation, which struggles to adapt to the dynamically varying information density in language generation, often failing to balance creativity and logical coherence. This work proposes Top-b decoding, the first approach to formalize the decoding process as an entropy-driven dynamic bandwidth modulation mechanism over a relative probability manifold. By coupling an adaptive bandwidth coefficient strictly with the model’s instantaneous Shannon entropy, Top-b dynamically adjusts the size of the candidate token set. Theoretical analysis demonstrates that Top-b minimizes tail distribution variance, thereby establishing an approximately self-regulating generative control system. Empirical results show that this method significantly reduces the variance between generation entropy and decoding behavior on benchmarks such as GPQA and GSM8K, while maintaining strong reasoning accuracy.

0 citationsRead paper