Enhanced Arabic-language cyberbullying detection: deep embedding and transformer (BERT) approaches

📅 2025-06-01
🏛️ IAES International Journal of Artificial Intelligence (IJ-AI)
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Research on Arabic cyberbullying detection remains scarce, hindered by the scarcity of annotated data and inherent challenges in modeling low-resource languages. Method: This work introduces a high-quality, manually annotated dataset comprising 10,662 social media posts and proposes a hybrid architecture integrating BERT with LSTM/Bi-LSTM to enhance semantic representation of Arabic text. Annotation quality is rigorously validated via Cohen’s Kappa. Multiple ablation studies compare FastText embeddings, pre-trained BERT, and their combinations. Contribution/Results: The Bi-LSTM-BERT model achieves 97% accuracy, while Bi-LSTM augmented with FastText attains 98%, significantly outperforming baseline models. Notably, both variants demonstrate strong cross-domain generalization. This study establishes a reproducible, benchmark-quality dataset and an effective, transferable modeling framework for cyberbullying detection in low-resource languages—particularly Arabic.

Technology Category

Application Category

📝 Abstract
Recent technological advances in smartphones and communications, including the growth of such online platforms as massive social media networks such as X (formerly known as Twitter) endangers young people and their emotional well-being by exposing them to cyberbullying, taunting, and bullying content. Most proposed approaches for automatically detecting cyberbullying have been developed around the English language, and methods for detecting Arabic-language cyberbullying are scarce. Methods for detecting Arabic-language cyberbullying are especially scarce. This paper aims to enhance the effectiveness of methods for detecting cyberbullying in Arabic-language content. We assembled a dataset of 10,662 X posts, pre-processed the data, and used the kappa tool to verify and enhance the quality of our annotations. We conducted four experiments to test numerous deep learning models for automatically detecting Arabic-language cyberbullying. We first tested a long short-term memory (LSTM) model and a bidirectional long short-term memory (Bi-LSTM) model with several experimental word embeddings. We also tested the LSTM and Bi-LSTM models with a novel pre-trained bidirectional encoder from representations (BERT) and then tested them on a different experimental models BERT again. LSTM-BERT and Bi-LSTM-BERT demonstrated a 97% accuracy. Bi-LSTM with FastText embedding word performed even better, achieving 98% accuracy. As a result, the outcomes are generalized.
Problem

Research questions and friction points this paper is trying to address.

Detecting cyberbullying in Arabic-language social media content
Addressing scarcity of effective Arabic cyberbullying detection methods
Evaluating deep learning models for Arabic cyberbullying identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Used BERT transformer for Arabic cyberbullying detection
Applied Bi-LSTM with FastText embedding achieving 98% accuracy
Combined LSTM-BERT models reaching 97% detection accuracy
🔎 Similar Papers
No similar papers found.
E
Ebtesam Jaber Aljohani
Department of Computer Science, College of Computer Science and Engineering, Taibah University, Madinah, Saudi Arabia
W
Wael M. S. Yafooz
Department of Computer Science, College of Computer Science and Engineering, Taibah University, Madinah, Saudi Arabia