The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了波斯语文本NLP中的标注瓶颈问题,通过对比分析34个资源并提出定量交叉检查方法来解决。
📝 Abstract
Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.
Problem

Research questions and friction points this paper is trying to address.

annotation-scarce
low-resource language
Persian NLP
resource ecology
task and domain coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

annotation-scarce
cross-checks
resource ecology
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
MohammadHossein Mortazavi
Network Science and Technology Group, School of Intelligent Systems, College of Interdisciplinary Science and Technologies, University of Tehran, Tehran, Iran
Mostafa Salehi
Mostafa Salehi
Associate Professor, University of Tehran
Social Network and Media AnalysisNetwork Science
Hadi Veisi
Hadi Veisi
Associate Professor, University of Tehran
Speech ProcessingNatural Language ProcessingArtificial Neural Network and Deep LearningMachine