Sample-Size Scaling of the African Languages NLI Evaluation

πŸ“… 2026-06-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the scarcity of annotated data for natural language inference (NLI) in African languages and investigates whether increased data volume consistently improves performance. Conducting controlled scaling experiments on the AfriXNLI benchmark, the authors evaluate XLM-R Large and AfroXLM-R Large across 16 African languages using sample sizes ranging from 50 to 500, with multiple random subsampling runs to assess robustness. The findings reveal a non-monotonic, highly language-dependent relationship between sample size and NLI performance: for some languages, performance plateaus or even declines with more data, and variance remains substantial in low-resource settings. These results challenge the conventional assumption that β€œmore data is always better,” suggesting that merely expanding labeled datasets may not reliably enhance NLI performance for African languages.
πŸ“ Abstract
African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance. The study is a systematic sample-size scaling study of natural language inference (NLI) on 16 African languages based on the AfriXNLI benchmark. Under controlled conditions, two multilingual transformer models with roughly 0.6B parameters XLM-R Large fine-tuned on XNLI and AfroXLM-R Large are tested on sample sizes of between 50 and 500 labeled examples and average their results across random subsampling runs. As opposed to the usual belief of monotonic increase with increased data, we find a strongly language sensitive and often non-monotonic scaling behavior. Some languages show early saturation or decrease in performance with sample size as well as high variance in low resource regimes. These results indicate that the volume of data is not enough to guarantee stable profits to African NLI, creating the necessity of language sensitive datasets creation and stronger multi-lingual modelling strategies.
Problem

Research questions and friction points this paper is trying to address.

African languages
natural language inference
sample-size scaling
low-resource NLP
data efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

sample-size scaling
African languages
natural language inference
low-resource NLP
multilingual modeling
πŸ’Ό Related Jobs
No related jobs found.
A
Anuj Tiwari
Noida Institute of Engineering and Technology
O
Oluwapelumi Ogunremu
ML Collective
T
Terry Oko-odion
ML Collective
J
Jesujuwon Egbewale
ML Collective
H
Hannah Nwokocha
ML Collective