Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection

📅 2026-07-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of Bangla hate speech detection models in real-world scenarios, particularly their poor handling of implicit expressions, cultural context, and emoji semantics, which often leads to misclassification and over-censorship. The authors systematically evaluate six mainstream architectures—including FastText, CNN/LSTM/BiLSTM variants, and BanglaBERT—on both benchmark and multi-source merged datasets, and for the first time expose their generalization failure using external data from real social platforms. They propose an emoji-aware preprocessing strategy that preserves and enhances emoji semantics, achieving up to a 12% F1 improvement on implicit hate speech detection. Results show BanglaBERT attains 91.4% F1 on benchmark data but drops sharply to 63.4% on external samples containing sarcasm and emojis, underscoring the need for culturally sensitive, context-adaptive ethical moderation frameworks.
📝 Abstract
The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced languages like Bangla, due to cultural context, implicit expressions, and informal linguistic patterns. This study aimed to expose the crisis of Bangla HS detection systems by diagnosing how and why benchmark-trained models fail to identify implicit, context-dependent HS. Six architectures (FastText + CNN, FastText + LSTM, FastText + BiLSTM, BanglaBERT, BanglaBERT + CNN, and BanglaBERT + BiLSTM) were trained on benchmark datasets (about 75,000 posts) and a merged multi-source dataset (about 120,000 posts), then externally validated on an annotated dataset (about 200 posts) collected from Facebook, Twitter, and YouTube, labeled as HS and non-HS, where HS was further categorized as explicit and implicit. BanglaBERT achieved an F1-score of 91.4% on benchmark datasets but declined to 75.3% on the external set and 63.4% for implicit HS involving sarcasm and emojis. The accuracy of FastText + CNN dropped from 78.0% to 51.2% under similar conditions. Emoji-aware preprocessing improved implicit HS detection by up to 12%, whereas emoji removal caused a notable decline in performance (F1: 0.75 to 0.63). Frequent misclassifications in politically charged or satirical comments revealed over-policing risks. This study not only exposes the generalization crisis due to implicit, culturally embedded, and emoji-laden expressions but also underscores the need for developing adaptive, emoji-aware, and culturally grounded frameworks that ensure ethical moderation while preserving freedom of expression. Findings of this study provide insights for researchers, SMPs, and policymakers to design more context-sensitive HS detection systems for low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

hate speech detection
implicit hate speech
Bangla language
emoji
generalization crisis
Innovation

Methods, ideas, or system contributions that make the work stand out.

hate speech detection
implicit hate speech
emoji-aware preprocessing
cross-dataset generalization
low-resource languages
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Faria Afrin Tisha
Department of Electrical and Computer Engineering, Rajshahi University of Engineering and Technology, Rajshahi, 6204, Bangladesh
Fariya Tabassum
Fariya Tabassum
Assistant Professor, Rajshahi University of Engineering & Technology
Power system
Hafsa Binte Kibria
Hafsa Binte Kibria
Assistant Professor, Department of Electrical & Computer Engineering, RUET
Machine LearningData MiningAI
M
Md. Nahiduzzaman
Department of Electrical and Computer Engineering, Rajshahi University of Engineering and Technology, Rajshahi, 6204, Bangladesh
Mominul Ahsan
Mominul Ahsan
Associate Lecturer
PrognosticsArtificial IntelligenceMachine learningData and Image analysisPower electronics