🤖 AI Summary
This study addresses the scarcity of publicly available datasets for evaluating cross-target generalization in Arabic stance detection. To bridge this gap, the authors introduce Mawqif-v2, an expanded dataset comprising 996 manually annotated Arabic tweets spanning three distinct topics: women driving, electric vehicles, and the trimester academic system. This work presents the first benchmark specifically designed for cross-target stance detection in Arabic, with each tweet labeled for stance, sentiment, and sarcasm. The dataset enables systematic evaluation of both zero-shot large language models and Arabic/multilingual Transformer-based approaches. By establishing reproducible baseline performance metrics, the study provides a standardized evaluation framework to advance research on cross-target generalization in Arabic stance detection.
📝 Abstract
Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-v2 Extension, consisting of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is annotated with stance, sentiment, and sarcasm labels following the original Mawqif annotation scheme. The released extension is intended as a held-out evaluation set for assessing model generalization to both semantically related and previously unseen targets, while the original Mawqif dataset is used for training and development. In addition, we establish baseline results using several Arabic and multilingual transformer models, as well as zero-shot large language models (LLMs), to facilitate reproducible evaluation. Together with the original Mawqif dataset, the Mawqif-v2 Extension provides a benchmark for evaluating cross-target generalization in Arabic stance detection.