Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of a fully functional electronic explanatory dictionary for Tajik and the inadequate adaptation of natural language processing (NLP) methods to this low-resource language. To bridge this gap, the authors propose a unified architecture that integrates traditional lexicography, linguistic statistics, and the generative capabilities of large language models (LLMs). The framework incorporates morphological analysis, tokenization, and semantic clustering, alongside subword segmentation and parameter-efficient fine-tuning (PEFT) strategies tailored for low-resource settings. This work presents the first systematic explanatory dictionary framework specifically designed for Tajik and establishes both a methodological foundation and technical support for downstream NLP applications such as machine translation, automatic summarization, and sentiment analysis.
📝 Abstract
This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lexicographic resource for Tajik that is comparable in functionality to dictionaries for high-resource languages, and from the limited adaptation of modern natural language processing technologies to low-resource language systems. Based on a systematic survey of existing linguistic, statistical, and corpus resources, we propose a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using LLMs. The choice of subword tokenization is justified by the agglutinative nature of Tajik morphology and its high morphological variability, along with a parameter-efficient fine-tuning (PEFT) strategy suitable for limited annotated data. The novelty of the work lies in proposing the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system. The practical significance of the study is the formation of a methodological foundation for developing a full-featured electronic dictionary that can serve both as a lexicographic tool and as a core resource for machine translation, automatic summarization, sentiment analysis, and other applied NLP tasks. The paper is intended for specialists in computational linguistics, lexicography, and developers of natural language processing systems working with low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

low-resource languages
Tajik language
electronic explanatory dictionary
natural language processing
lexicographic resource
Innovation

Methods, ideas, or system contributions that make the work stand out.

low-resource languages
large language models
electronic explanatory dictionary
parameter-efficient fine-tuning
morphological analysis
M
Mullosharaf K. Arabov
Kazan Federal University, Institute of Computational Mathematics and Information Technologies, Kazan, Russia