ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
构建了ParsHate数据集,用于波斯语中的仇恨言论和目标检测,通过十年间10,000条推文标注,支持多种分类任务,挑战现有模型性能。
📝 Abstract
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Hate Speech Detection
Persian
Target Identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark Dataset
Hate Speech Detection
Multi-label Target Identification
Persian Language
Span-level Rationales
🔎 Similar Papers
No similar papers found.