Human-Annotated NER Dataset for the Kyrgyz Language
To address the scarcity of high-quality manually annotated data for Kyrgyz named entity recognition (NER), this study introduces KyrgyzNER—the first large-scale, fine-grained, human-annotated NER dataset for Kyrgyz, comprising 1,499 news articles, 10,900 sentences, and 39,075 entity mentions across 27 entity types. We establish linguistically informed annotation guidelines tailored to Kyrgyz morphology and syntax. We systematically evaluate both traditional conditional random fields (CRF) and multilingual pre-trained language models (e.g., mRoBERTa) on this benchmark. Experimental results demonstrate that mRoBERTa achieves the best trade-off between precision and recall, significantly outperforming CRF-based approaches; other multilingual models also exhibit robust performance. KyrgyzNER fills a critical gap in NER resources for low-resource Turkic languages and serves as a foundational benchmark for cross-lingual information extraction research, providing both empirical evidence and standardized evaluation infrastructure.