E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决阿拉伯语在自然语言推理资源有限的问题,本文通过多种来源构建E-CONAN基准数据集,并用其评估了多语言预训练模型和特定于阿拉伯语的模型。
📝 Abstract
Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all around the world, Arabic language still suffers from limited resources in this domain. To address this gap, this paper introduces E-CONAN benchmarks that are composed of sentences pairs from various sources: (1) automatically-translated pairs, (2) human-validated machine-translated pairs, (3) hand-crafted pairs from teaching Arabic as foreign language books, and (4) headlines pairs from different news channels containing rumors. E-CONAN contains two benchmark datasets, E-CONAN-2, a 2-way dataset (RTE) and E-CONAN-3, a 3-way dataset (NLI). Additionally, we have used E-CONAN benchmarks to evaluate 9 state-of-the-art multilingual pretrained models using zero-shot classification. Models were evaluated across the ArNLI, XNLI, and E-CONAN datasets. Results show that E-CONAN is a potentially valuable resource for evaluating model generalization and even for fine-tuning pre-trained models. Its diverse composition, derived from a combination of sources, offers a broader and more robust assessment compared to XNLI and ArNLI. In addition, we have evaluated 5 LLMs on E-CONAN-3 dataset. Moreover, we incorporated MARBERT as a representative Arabic-specific baseline and conducted performance evaluation comparison to demonstrate how Arabic-specific models scale against cross-lingual and LLM-based approaches on the E-CONAN benchmarks. Furthermore, we conducted detailed qualitative and quantitative error analysis to analyze frequent error patterns. E-CONAN benchmarks will be publicly available, we hope that it will enrich research community in Arabic textual entailment and natural language inference.
Problem

Research questions and friction points this paper is trying to address.

Arabic
Natural Language Inference
Textual Entailment
Benchmark
Datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

E-CONAN
Arabic Textual Entailment
Natural Language Inference
Diverse Datasets
Zero-shot Classification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Khloud AL Jallad
Higher Institute for Applied Sciences and Technology, Damascus, Syria
Nada Ghneim
Nada Ghneim
Arab International University
Engineering & Technology / Computer Science
G
Ghaida Rebdawi
Higher Institute for Applied Sciences and Technology, Damascus, Syria