A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
该研究通过定量元分析,使用BERTopic等方法对7,120篇阿拉伯语NLP论文进行主题建模和引用预测,识别出研究趋势、主题演变及研究空白。
该研究通过定量元分析,使用BERTopic等方法对7,120篇阿拉伯语NLP论文进行主题建模和引用预测,识别出研究趋势、主题演变及研究空白。
This study addresses the lack of a fully functional electronic explanatory dictionary for Tajik and the inadequate adaptation of natural language processing (NLP) methods to this low-resource language. To bridge this gap, the authors propose a unified architecture that integrates traditional lexicography, linguistic statistics, and the generative capabilities of large language models (LLMs). The framework incorporates morphological analysis, tokenization, and semantic clustering, alongside subword segmentation and parameter-efficient fine-tuning (PEFT) strategies tailored for low-resource settings. This work presents the first systematic explanatory dictionary framework specifically designed for Tajik and establishes both a methodological foundation and technical support for downstream NLP applications such as machine translation, automatic summarization, and sentiment analysis.
This work addresses the absence of verifiable intermediate reasoning steps in existing Russian-language financial reasoning benchmarks, which are often limited to English or multiple-choice formats. The authors propose RusFinChain, the first Russian financial symbolic reasoning benchmark, spanning 17 domains and 172 topics, comprising 5,280 parametric samples generated via executable Python templates. Each sample includes a gold reasoning chain with intermediate numerical values, ensuring data contamination isolation and enabling automatic verification. The study introduces novel multidimensional evaluation metrics—such as fuzzy numerical alignment and soft attention alignment—that substantially improve correlation with final answer correctness (Spearman’s ρ = 0.48). Experiments on eight open-source large language models reveal a stark gap: while step-level alignment (Hard F1) reaches 0.65, answer accuracy remains around 29%, highlighting current models’ deficiencies in rigorous multi-step financial reasoning.
This study addresses the language barrier impeding knowledge exchange between Arabic- and Russian-speaking scientific communities, which hinders collaboration toward Sustainable Development Goals 9 and 17. To bridge this gap, the authors construct the first Arabic–Russian scientific parallel corpus, comprising approximately 27,000 sentence pairs, and perform efficient domain-specific fine-tuning of multilingual large language models—including mT5-base, NLLB, and Qwen2.5-7B-Instruct—using LoRA and QLoRA. Experimental results demonstrate that fine-tuned models significantly outperform few-shot prompting; notably, Qwen2.5-7B-Instruct with QLoRA (rank=8) achieves a BLEU score of 23.15 and a COMET score of 0.758. The project releases the corpus, fine-tuned models, and evaluation code, establishing the first benchmark for scientific translation between these two languages.
This study addresses the challenges of modeling and controlling the pectin hydrolysis-extraction process, which involves highly coupled parameters and extensive experimental dependency. Leveraging a dataset of 1,000 experiments, it presents the first systematic comparison of 11 machine learning algorithms for simultaneously predicting four critical quality indicators: pectin yield, galacturonic acid content, molecular weight, and degree of esterification. The CatBoost model achieves the best performance (mean R² = 0.946), with raw material type identified as the most influential factor, contributing 63.6% to predictive accuracy. By integrating explainable AI techniques and hyperparameter optimization, the authors develop an end-to-end intelligent modeling and deployment pipeline, culminating in an interactive web-based control system that substantially reduces experimental costs and advances the intelligent manufacturing of pectin.
该研究通过定量元分析,使用BERTopic等方法对7,120篇阿拉伯语NLP论文进行主题建模和引用预测,识别出研究趋势、主题演变及研究空白。
This study addresses the lack of a fully functional electronic explanatory dictionary for Tajik and the inadequate adaptation of natural language processing (NLP) methods to this low-resource language. To bridge this gap, the authors propose a unified architecture that integrates traditional lexicography, linguistic statistics, and the generative capabilities of large language models (LLMs). The framework incorporates morphological analysis, tokenization, and semantic clustering, alongside subword segmentation and parameter-efficient fine-tuning (PEFT) strategies tailored for low-resource settings. This work presents the first systematic explanatory dictionary framework specifically designed for Tajik and establishes both a methodological foundation and technical support for downstream NLP applications such as machine translation, automatic summarization, and sentiment analysis.
This work addresses the absence of verifiable intermediate reasoning steps in existing Russian-language financial reasoning benchmarks, which are often limited to English or multiple-choice formats. The authors propose RusFinChain, the first Russian financial symbolic reasoning benchmark, spanning 17 domains and 172 topics, comprising 5,280 parametric samples generated via executable Python templates. Each sample includes a gold reasoning chain with intermediate numerical values, ensuring data contamination isolation and enabling automatic verification. The study introduces novel multidimensional evaluation metrics—such as fuzzy numerical alignment and soft attention alignment—that substantially improve correlation with final answer correctness (Spearman’s ρ = 0.48). Experiments on eight open-source large language models reveal a stark gap: while step-level alignment (Hard F1) reaches 0.65, answer accuracy remains around 29%, highlighting current models’ deficiencies in rigorous multi-step financial reasoning.
This study addresses the language barrier impeding knowledge exchange between Arabic- and Russian-speaking scientific communities, which hinders collaboration toward Sustainable Development Goals 9 and 17. To bridge this gap, the authors construct the first Arabic–Russian scientific parallel corpus, comprising approximately 27,000 sentence pairs, and perform efficient domain-specific fine-tuning of multilingual large language models—including mT5-base, NLLB, and Qwen2.5-7B-Instruct—using LoRA and QLoRA. Experimental results demonstrate that fine-tuned models significantly outperform few-shot prompting; notably, Qwen2.5-7B-Instruct with QLoRA (rank=8) achieves a BLEU score of 23.15 and a COMET score of 0.758. The project releases the corpus, fine-tuned models, and evaluation code, establishing the first benchmark for scientific translation between these two languages.
This study addresses the challenges of modeling and controlling the pectin hydrolysis-extraction process, which involves highly coupled parameters and extensive experimental dependency. Leveraging a dataset of 1,000 experiments, it presents the first systematic comparison of 11 machine learning algorithms for simultaneously predicting four critical quality indicators: pectin yield, galacturonic acid content, molecular weight, and degree of esterification. The CatBoost model achieves the best performance (mean R² = 0.946), with raw material type identified as the most influential factor, contributing 63.6% to predictive accuracy. By integrating explainable AI techniques and hyperparameter optimization, the authors develop an end-to-end intelligent modeling and deployment pipeline, culminating in an interactive web-based control system that substantially reduces experimental costs and advances the intelligent manufacturing of pectin.