MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection

๐Ÿ“… 2025-02-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Arabic, the fourth most widely used language online, suffers from a longstanding scarcity of multi-task annotated resources for propaganda, sentiment, and emotion analysis. To address this gap, we introduce MultiProSEโ€”the first Arabic multilabel dataset comprising 8,000 manually annotated news articles, jointly labeled for propaganda detection, sentiment classification, and fine-grained emotion identification. We extend the ArPro benchmark to establish an open-source multilingual multistage evaluation framework. Leveraging GPT-4o-mini and three BERT variants, we develop reproducible multistage baselines. We publicly release annotation guidelines, source code, and baseline results. This work fills a critical void in Arabic fine-grained opinion mining resources and provides foundational infrastructure for non-English propaganda detection and cross-dimensional stance understanding.

Technology Category

Application Category

๐Ÿ“ Abstract
Propaganda is a form of persuasion that has been used throughout history with the intention goal of influencing people's opinions through rhetorical and psychological persuasion techniques for determined ends. Although Arabic ranked as the fourth most- used language on the internet, resources for propaganda detection in languages other than English, especially Arabic, remain extremely limited. To address this gap, the first Arabic dataset for Multi-label Propaganda, Sentiment, and Emotion (MultiProSE) has been introduced. MultiProSE is an open-source extension of the existing Arabic propaganda dataset, ArPro, with the addition of sentiment and emotion annotations for each text. This dataset comprises 8,000 annotated news articles, which is the largest propaganda dataset to date. For each task, several baselines have been developed using large language models (LLMs), such as GPT-4o-mini, and pre-trained language models (PLMs), including three BERT-based models. The dataset, annotation guidelines, and source code are all publicly released to facilitate future research and development in Arabic language models and contribute to a deeper understanding of how various opinion dimensions interact in news media1.
Problem

Research questions and friction points this paper is trying to address.

Detects propaganda in Arabic news
Analyzes sentiment and emotion
Expands resources for Arabic NLP
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-label Arabic dataset
GPT-4o-mini and BERT models
8,000 annotated news articles
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
L
Lubna Al-Henaki
Majmaah University, King Saud University
H
H. Al-Khalifa
King Saud University
A
A. Al-Salman
King Saud University
H
Hajar Alqubayshi
Imam Mohammad Ibn Saud Islamic University
H
Hind Al-Twailay
Imam Mohammad Ibn Saud Islamic University
G
Gheeda Alghamdi
Alfaisal University
H
Hawra Aljasim
King Saud University