Institution profile

Kazan Federal University

Academic institutioneurope · ru
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

Aug 04, 2026

This study addresses the lack of a fully functional electronic explanatory dictionary for Tajik and the inadequate adaptation of natural language processing (NLP) methods to this low-resource language. To bridge this gap, the authors propose a unified architecture that integrates traditional lexicography, linguistic statistics, and the generative capabilities of large language models (LLMs). The framework incorporates morphological analysis, tokenization, and semantic clustering, alongside subword segmentation and parameter-efficient fine-tuning (PEFT) strategies tailored for low-resource settings. This work presents the first systematic explanatory dictionary framework specifically designed for Tajik and establishes both a methodological foundation and technical support for downstream NLP applications such as machine translation, automatic summarization, and sentiment analysis.

0 citationsRead paper

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Jul 01, 2026

This work addresses the absence of verifiable intermediate reasoning steps in existing Russian-language financial reasoning benchmarks, which are often limited to English or multiple-choice formats. The authors propose RusFinChain, the first Russian financial symbolic reasoning benchmark, spanning 17 domains and 172 topics, comprising 5,280 parametric samples generated via executable Python templates. Each sample includes a gold reasoning chain with intermediate numerical values, ensuring data contamination isolation and enabling automatic verification. The study introduces novel multidimensional evaluation metrics—such as fuzzy numerical alignment and soft attention alignment—that substantially improve correlation with final answer correctness (Spearman’s ρ = 0.48). Experiments on eight open-source large language models reveal a stark gap: while step-level alignment (Hard F1) reaches 0.65, answer accuracy remains around 29%, highlighting current models’ deficiencies in rigorous multi-step financial reasoning.

0 citationsRead paper

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

Jun 29, 2026

This study addresses the language barrier impeding knowledge exchange between Arabic- and Russian-speaking scientific communities, which hinders collaboration toward Sustainable Development Goals 9 and 17. To bridge this gap, the authors construct the first Arabic–Russian scientific parallel corpus, comprising approximately 27,000 sentence pairs, and perform efficient domain-specific fine-tuning of multilingual large language models—including mT5-base, NLLB, and Qwen2.5-7B-Instruct—using LoRA and QLoRA. Experimental results demonstrate that fine-tuned models significantly outperform few-shot prompting; notably, Qwen2.5-7B-Instruct with QLoRA (rank=8) achieves a BLEU score of 23.15 and a COMET score of 0.758. The project releases the corpus, fine-tuned models, and evaluation code, establishing the first benchmark for scientific translation between these two languages.

0 citationsRead paper

A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

May 30, 2026

This study addresses the challenges of modeling and controlling the pectin hydrolysis-extraction process, which involves highly coupled parameters and extensive experimental dependency. Leveraging a dataset of 1,000 experiments, it presents the first systematic comparison of 11 machine learning algorithms for simultaneously predicting four critical quality indicators: pectin yield, galacturonic acid content, molecular weight, and degree of esterification. The CatBoost model achieves the best performance (mean R² = 0.946), with raw material type identified as the most influential factor, contributing 63.6% to predictive accuracy. By integrating explainable AI techniques and hyperparameter optimization, the authors develop an end-to-end intelligent modeling and deployment pipeline, culminating in an interactive web-based control system that substantially reduces experimental costs and advances the intelligent manufacturing of pectin.

0 citationsRead paper
Recent publications

Latest Papers

Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

Aug 04, 2026

This study addresses the lack of a fully functional electronic explanatory dictionary for Tajik and the inadequate adaptation of natural language processing (NLP) methods to this low-resource language. To bridge this gap, the authors propose a unified architecture that integrates traditional lexicography, linguistic statistics, and the generative capabilities of large language models (LLMs). The framework incorporates morphological analysis, tokenization, and semantic clustering, alongside subword segmentation and parameter-efficient fine-tuning (PEFT) strategies tailored for low-resource settings. This work presents the first systematic explanatory dictionary framework specifically designed for Tajik and establishes both a methodological foundation and technical support for downstream NLP applications such as machine translation, automatic summarization, and sentiment analysis.

0 citationsRead paper

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Jul 01, 2026

This work addresses the absence of verifiable intermediate reasoning steps in existing Russian-language financial reasoning benchmarks, which are often limited to English or multiple-choice formats. The authors propose RusFinChain, the first Russian financial symbolic reasoning benchmark, spanning 17 domains and 172 topics, comprising 5,280 parametric samples generated via executable Python templates. Each sample includes a gold reasoning chain with intermediate numerical values, ensuring data contamination isolation and enabling automatic verification. The study introduces novel multidimensional evaluation metrics—such as fuzzy numerical alignment and soft attention alignment—that substantially improve correlation with final answer correctness (Spearman’s ρ = 0.48). Experiments on eight open-source large language models reveal a stark gap: while step-level alignment (Hard F1) reaches 0.65, answer accuracy remains around 29%, highlighting current models’ deficiencies in rigorous multi-step financial reasoning.

0 citationsRead paper

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

Jun 29, 2026

This study addresses the language barrier impeding knowledge exchange between Arabic- and Russian-speaking scientific communities, which hinders collaboration toward Sustainable Development Goals 9 and 17. To bridge this gap, the authors construct the first Arabic–Russian scientific parallel corpus, comprising approximately 27,000 sentence pairs, and perform efficient domain-specific fine-tuning of multilingual large language models—including mT5-base, NLLB, and Qwen2.5-7B-Instruct—using LoRA and QLoRA. Experimental results demonstrate that fine-tuned models significantly outperform few-shot prompting; notably, Qwen2.5-7B-Instruct with QLoRA (rank=8) achieves a BLEU score of 23.15 and a COMET score of 0.758. The project releases the corpus, fine-tuned models, and evaluation code, establishing the first benchmark for scientific translation between these two languages.

0 citationsRead paper

A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

May 30, 2026

This study addresses the challenges of modeling and controlling the pectin hydrolysis-extraction process, which involves highly coupled parameters and extensive experimental dependency. Leveraging a dataset of 1,000 experiments, it presents the first systematic comparison of 11 machine learning algorithms for simultaneously predicting four critical quality indicators: pectin yield, galacturonic acid content, molecular weight, and degree of esterification. The CatBoost model achieves the best performance (mean R² = 0.946), with raw material type identified as the most influential factor, contributing 63.6% to predictive accuracy. By integrating explainable AI techniques and hyperparameter optimization, the authors develop an end-to-end intelligent modeling and deployment pipeline, culminating in an interactive web-based control system that substantially reduces experimental costs and advances the intelligent manufacturing of pectin.

0 citationsRead paper