Institution profile

PDMI RAS

Academic institutioneurope · ru
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Human-Annotated NER Dataset for the Kyrgyz Language

Sep 23, 2025

To address the scarcity of high-quality manually annotated data for Kyrgyz named entity recognition (NER), this study introduces KyrgyzNER—the first large-scale, fine-grained, human-annotated NER dataset for Kyrgyz, comprising 1,499 news articles, 10,900 sentences, and 39,075 entity mentions across 27 entity types. We establish linguistically informed annotation guidelines tailored to Kyrgyz morphology and syntax. We systematically evaluate both traditional conditional random fields (CRF) and multilingual pre-trained language models (e.g., mRoBERTa) on this benchmark. Experimental results demonstrate that mRoBERTa achieves the best trade-off between precision and recall, significantly outperforming CRF-based approaches; other multilingual models also exhibit robust performance. KyrgyzNER fills a critical gap in NER resources for low-resource Turkic languages and serves as a foundational benchmark for cross-lingual information extraction research, providing both empirical evidence and standardized evaluation infrastructure.

0 citationsRead paper
Recent publications

Latest Papers

Human-Annotated NER Dataset for the Kyrgyz Language

Sep 23, 2025

To address the scarcity of high-quality manually annotated data for Kyrgyz named entity recognition (NER), this study introduces KyrgyzNER—the first large-scale, fine-grained, human-annotated NER dataset for Kyrgyz, comprising 1,499 news articles, 10,900 sentences, and 39,075 entity mentions across 27 entity types. We establish linguistically informed annotation guidelines tailored to Kyrgyz morphology and syntax. We systematically evaluate both traditional conditional random fields (CRF) and multilingual pre-trained language models (e.g., mRoBERTa) on this benchmark. Experimental results demonstrate that mRoBERTa achieves the best trade-off between precision and recall, significantly outperforming CRF-based approaches; other multilingual models also exhibit robust performance. KyrgyzNER fills a critical gap in NER resources for low-resource Turkic languages and serves as a foundational benchmark for cross-lingual information extraction research, providing both empirical evidence and standardized evaluation infrastructure.

0 citationsRead paper