SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical gap in safety evaluations of large language models, which have predominantly focused on English and Western contexts while overlooking the cultural diversity and local sensitivities of multilingual regions such as India. The work proposes the first systematically constructed, human-authored safety benchmark encompassing ten major Indian languages alongside English, integrating both general and region-specific prompts. It employs a fine-grained safety taxonomy and a context-aware evaluation protocol. Comprehensive assessments of prominent multilingual models reveal pervasive issues in Indian languages—particularly when written in native scripts—including excessive refusal rates, failure to detect implicit biases, and inadequate understanding of regional context. These findings underscore the urgent need for localized safety evaluation frameworks beyond English-centric paradigms.
📝 Abstract
Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.
Problem

Research questions and friction points this paper is trying to address.

safety evaluation
multilingual LLMs
Indic languages
cultural sensitivity
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

safety benchmark
multilingual LLMs
Indic languages
culturally grounded evaluation
region-specific prompts
💼 Related Jobs
No related jobs found.