EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)

📅 2025-04-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of domain-specific annotated data and the difficulty in identifying education-related entities—such as academic roles, course names, and institutional terms—in Urdu educational texts, introducing for the first time the Urdu educational-domain named entity recognition (NER) task. We construct EDU-NER-2025, the first manually annotated, domain-specific NER dataset for Urdu education, comprising 13 fine-grained entity types. To tackle morphological complexity and semantic ambiguity, we perform domain-adaptive fine-tuning of XLM-RoBERTa and augment training data with real-world educational tweets. Our model achieves an F1 score of 86.3% on the EDU-NER-2025 test set, significantly outperforming baseline models. To foster community advancement, we publicly release the dataset, annotation guidelines, and implementation code—thereby filling a critical gap in Urdu educational NLP resources.

Technology Category

Application Category

📝 Abstract
Named Entity Recognition (NER) plays a pivotal role in various Natural Language Processing (NLP) tasks by identifying and classifying named entities (NEs) from unstructured data into predefined categories such as person, organization, location, date, and time. While extensive research exists for high-resource languages and general domains, NER in Urdu particularly within domain-specific contexts like education remains significantly underexplored. This is Due to lack of annotated datasets for educational content which limits the ability of existing models to accurately identify entities such as academic roles, course names, and institutional terms, underscoring the urgent need for targeted resources in this domain. To the best of our knowledge, no dataset exists in the domain of the Urdu language for this purpose. To achieve this objective this study makes three key contributions. Firstly, we created a manually annotated dataset in the education domain, named EDU-NER-2025, which contains 13 unique most important entities related to education domain. Second, we describe our annotation process and guidelines in detail and discuss the challenges of labelling EDU-NER-2025 dataset. Third, we addressed and analyzed key linguistic challenges, such as morphological complexity and ambiguity, which are prevalent in formal Urdu texts.
Problem

Research questions and friction points this paper is trying to address.

Lack of Urdu NER datasets for educational texts
Difficulty in recognizing domain-specific entities in Urdu
Challenges in Urdu text annotation due to linguistic complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Used XLM-RoBERTa for Urdu NER
Created EDU-NER-2025 annotated dataset
Addressed Urdu morphological complexity challenges
🔎 Similar Papers
No similar papers found.
Fida Ullah
Fida Ullah
PhD Computer Science
Natural Language ProcessingNamed Entity RecognitionDeep LearningComputational Linguistic.
Muhammad Ahmad
Muhammad Ahmad
King Fahd University of Petroleum and Minerals
Machine LearningComputer VisionHyperspectral imaging
M
M. Zamir
Centro de Investigación en Computación, Instituto Politécnico Nacional (CIC -PN), Mexico City 07738, Mexico
M
Muhammad Arif
Centro de Investigación en Computación, Instituto Politécnico Nacional (CIC -PN), Mexico City 07738, Mexico
Grigori Sidorov
Grigori Sidorov
Professor of Computational Linguistics, Instituto Politécnico Nacional (IPN), Mexico
Computational LinguisticsNatural Language ProcessingArtificial IntelligenceMachine Learning
E
Edgardo Manuel Felipe River'on
Centro de Investigación en Computación, Instituto Politécnico Nacional (CIC -PN), Mexico City 07738, Mexico
A
Alexander F. Gelbukh
Centro de Investigación en Computación, Instituto Politécnico Nacional (CIC -PN), Mexico City 07738, Mexico