Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了孟加拉语中LLM安全性评估不足的问题,通过创建包含879个提示的BanglaSafe基准,并发现写作风格对安全性的显著影响。
📝 Abstract
Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed prompts, spanning 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Evaluating 18 frontier LLMs, we find that over half of all responses are unsafe or partially unsafe (53.6%) while 14.7% contains strictly harmful content, and that the strongest observed effect is not the switch from English to Bengali but the choice of writing style within Bengali: the same harmful request phrased as a formal newspaper investigation succeeds 17 percentage points more often than the same request phrased as a casual message, with no adversarial engineering involved. We further show that existing safety classifiers struggle to reliably evaluate Bengali content, with even frontier models failing on nearly half of all cases.
Problem

Research questions and friction points this paper is trying to address.

LLM Safety
Bengali
Culturally Grounded Harms
Safety Classifiers
Innovation

Methods, ideas, or system contributions that make the work stand out.

BanglaSafe
writing style impact
cultural harms
language safety evaluation
Bengali LLM
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Naymul Islam
BanglaLLM
N
Nusrat Jahan Lia
Institute of Information Technology, University of Dhaka
Shubhashis Roy Dipta
Shubhashis Roy Dipta
University of Maryland, Baltimore County
Natural Language ProcessingReasoningMultimodal Understanding
S
Sabik Bin Sultan
Bangladesh Air Force Shaheen College Kurmitola
A
Abdullah Khan Zehady
Ciroos Inc.