Institution profile

Indira Gandhi Delhi Technical University for Women

Academic institutionasia · in
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Jul 28, 2026

This study addresses the limitations of current mainstream large language model (LLM) evaluation benchmarks, which are predominantly English- and Western-centric and thus inadequately assess model safety, fairness, and accuracy in India’s multilingual and multicultural context. Building upon the UK AISI’s Inspect AI platform, the authors introduce the first open-source evaluation framework tailored to India’s 22 official languages. The framework encompasses six dimensions: multilingual MMLU, localized bias testing (BharatBBQ), multi-turn jailbreak resistance, cultural knowledge assessment, and safety related to digital public infrastructure (DPI), alongside an LLM-as-judge automated scoring mechanism. Evaluations of five open-source models (8B–32B parameters) reveal that Sarvam-M 24B and Gemma 2 27B both achieve 80% on an Indian fairness index, with Sarvam-M excelling in cultural knowledge and DPI compliance. While all models uniformly reject harmful multilingual prompts (100% refusal rate), their DPI safety scores vary widely (20%–100%).

0 citationsRead paper

Fraud Detection System for Banking Transactions

Apr 09, 2026

This study addresses the challenges of dynamically evolving financial fraud and severe class imbalance in digital payment systems by conducting hypothesis-driven exploratory data analysis and feature engineering on the PaySim synthetic dataset, following the CRISP-DM methodology. To mitigate class imbalance, SMOTE oversampling is employed, and hyperparameter optimization is performed via GridSearchCV across multiple classifiers, including logistic regression, decision trees, random forests, and XGBoost. The resulting fraud detection framework achieves significantly enhanced detection performance while maintaining high scalability and robustness, thereby offering FinTech systems an efficient and reliable solution for real-time fraud prevention.

0 citationsRead paper

Code-Mix Sentiment Analysis on Hinglish Tweets

Jan 08, 2026arXiv.org

This study addresses the challenges posed by Hinglish—a prevalent Romanized Hindi–English code-mixed variety on Indian social media—whose spelling variations, slang, and out-of-vocabulary terms significantly degrade the performance of conventional monolingual NLP models in sentiment analysis, thereby undermining brand monitoring efforts. To tackle this, the work proposes a high-performance sentiment classification framework specifically designed for Hinglish tweets, which uniquely integrates fine-tuned multilingual BERT (mBERT) with subword tokenization to effectively handle the low-resource nature of code-mixed language. Evaluated on public benchmarks, the approach achieves state-of-the-art accuracy, offering both a deployable, production-ready tool for brand sentiment tracking and a new strong baseline for code-mixing NLP tasks.

0 citationsRead paper
Recent publications

Latest Papers

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Jul 28, 2026

This study addresses the limitations of current mainstream large language model (LLM) evaluation benchmarks, which are predominantly English- and Western-centric and thus inadequately assess model safety, fairness, and accuracy in India’s multilingual and multicultural context. Building upon the UK AISI’s Inspect AI platform, the authors introduce the first open-source evaluation framework tailored to India’s 22 official languages. The framework encompasses six dimensions: multilingual MMLU, localized bias testing (BharatBBQ), multi-turn jailbreak resistance, cultural knowledge assessment, and safety related to digital public infrastructure (DPI), alongside an LLM-as-judge automated scoring mechanism. Evaluations of five open-source models (8B–32B parameters) reveal that Sarvam-M 24B and Gemma 2 27B both achieve 80% on an Indian fairness index, with Sarvam-M excelling in cultural knowledge and DPI compliance. While all models uniformly reject harmful multilingual prompts (100% refusal rate), their DPI safety scores vary widely (20%–100%).

0 citationsRead paper

Fraud Detection System for Banking Transactions

Apr 09, 2026

This study addresses the challenges of dynamically evolving financial fraud and severe class imbalance in digital payment systems by conducting hypothesis-driven exploratory data analysis and feature engineering on the PaySim synthetic dataset, following the CRISP-DM methodology. To mitigate class imbalance, SMOTE oversampling is employed, and hyperparameter optimization is performed via GridSearchCV across multiple classifiers, including logistic regression, decision trees, random forests, and XGBoost. The resulting fraud detection framework achieves significantly enhanced detection performance while maintaining high scalability and robustness, thereby offering FinTech systems an efficient and reliable solution for real-time fraud prevention.

0 citationsRead paper

Code-Mix Sentiment Analysis on Hinglish Tweets

Jan 08, 2026arXiv.org

This study addresses the challenges posed by Hinglish—a prevalent Romanized Hindi–English code-mixed variety on Indian social media—whose spelling variations, slang, and out-of-vocabulary terms significantly degrade the performance of conventional monolingual NLP models in sentiment analysis, thereby undermining brand monitoring efforts. To tackle this, the work proposes a high-performance sentiment classification framework specifically designed for Hinglish tweets, which uniquely integrates fine-tuned multilingual BERT (mBERT) with subword tokenization to effectively handle the low-resource nature of code-mixed language. Evaluated on public benchmarks, the approach achieves state-of-the-art accuracy, offering both a deployable, production-ready tool for brand sentiment tracking and a new strong baseline for code-mixing NLP tasks.

0 citationsRead paper