Institution profile

KRUTRIM

Industry researchasia · in
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

Nov 13, 2025

To address the scarcity of pretraining data for low-resource Indian languages, this paper introduces BhashaKritika, a multilingual synthetic data construction framework covering 10 Indian languages and 54 billion tokens. Methodologically, it proposes the first document–role–topic co-guided synthetic generation paradigm, integrating five complementary generation techniques; it further designs a modular quality assurance pipeline incorporating script/language identification, metadata consistency verification, n-gram deduplication, and KenLM perplexity filtering—enabling efficient cross-script and cross-lingual quality control. Comprehensive experiments systematically characterize the quality–diversity trade-offs across generation strategies, establishing best practices for multilingual synthetic corpus construction. Empirical results demonstrate that models pretrained on BhashaKritika achieve substantial performance gains across Indian languages, providing a reusable data infrastructure and methodological blueprint for low-resource multilingual LLM development.

1 citationsRead paper

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

Jun 17, 2026

This study addresses the significant performance degradation of foundation automatic speech recognition (ASR) models under narrowband channel conditions and in low-resource settings involving languages such as Hindi and accents like Indian English. The authors systematically evaluate the zero-shot capabilities of prominent open-source and commercial foundation ASR models on real-world narrowband telephone speech and investigate the effectiveness of fine-tuning with limited labeled data. Results demonstrate that zero-shot models generally perform poorly, while supervised fine-tuning with small amounts of data yields performance gains that vary substantially across languages and accents and are strongly dependent on the scale of pretraining data. The work identifies critical bottlenecks of foundation ASR systems in low-resource scenarios and quantifies the cross-lingual generalization potential of fine-tuning strategies.

0 citationsRead paper
Recent publications

Latest Papers

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

Jun 17, 2026

This study addresses the significant performance degradation of foundation automatic speech recognition (ASR) models under narrowband channel conditions and in low-resource settings involving languages such as Hindi and accents like Indian English. The authors systematically evaluate the zero-shot capabilities of prominent open-source and commercial foundation ASR models on real-world narrowband telephone speech and investigate the effectiveness of fine-tuning with limited labeled data. Results demonstrate that zero-shot models generally perform poorly, while supervised fine-tuning with small amounts of data yields performance gains that vary substantially across languages and accents and are strongly dependent on the scale of pretraining data. The work identifies critical bottlenecks of foundation ASR systems in low-resource scenarios and quantifies the cross-lingual generalization potential of fine-tuning strategies.

0 citationsRead paper

BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages

Nov 13, 2025

To address the scarcity of pretraining data for low-resource Indian languages, this paper introduces BhashaKritika, a multilingual synthetic data construction framework covering 10 Indian languages and 54 billion tokens. Methodologically, it proposes the first document–role–topic co-guided synthetic generation paradigm, integrating five complementary generation techniques; it further designs a modular quality assurance pipeline incorporating script/language identification, metadata consistency verification, n-gram deduplication, and KenLM perplexity filtering—enabling efficient cross-script and cross-lingual quality control. Comprehensive experiments systematically characterize the quality–diversity trade-offs across generation strategies, establishing best practices for multilingual synthetic corpus construction. Empirical results demonstrate that models pretrained on BhashaKritika achieve substantial performance gains across Indian languages, providing a reusable data infrastructure and methodological blueprint for low-resource multilingual LLM development.

1 citationsRead paper