Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多语言嵌入模型评估数据集稀缺问题,提出一个两维框架,通过分析斯拉夫语系子集的任务特定和跨任务评估来解决。
📝 Abstract
Multilingual text embedding models enable cross-lingual transfer of knowledge across a wide range of NLP tasks, but their evaluation remains highly uneven across high-, mid- and low-resource languages. In this paper, we propose a two-dimensional framework, specifically tailored for analyzing multilingual embedding benchmarks under dataset scarcity, and apply it on the Slavic-language subset of the MTEB benchmark. The framework distinguishes between task-specific and cross-task evaluation, while jointly analyzing three complementary aspects: (1) ranking robustness, (2) model consistency, and (3) evidence strength. At the task-specific level, we evaluate the stability of model rankings under changes in ranking methodology and benchmark dataset composition. At the cross-task level, we assess the ability of models to generalize across diverse tasks within a language. To quantify the reliability of benchmark conclusions, we introduce an Evidence Strength Score that accounts for dataset availability, diversity, and robustness assessability. Our analysis reveals severe benchmark sparsity, with many Slavic language-task pairs relying on a single dataset or highly correlated benchmark collections, limiting the ability to draw robust conclusions. The cross-task analysis reveals a small group of highly transferable models, most notably llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants, that consistently perform well across Slavic languages and tasks. Overall, the results demonstrate that benchmark rankings and robustness conclusions must be interpreted jointly with certain notation of their evidence strength and highlight benchmark scarcity as a major obstacle to trustworthy multilingual evaluation.
Problem

Research questions and friction points this paper is trying to address.

multilingual embedding models
dataset scarcity
evaluation robustness
Slavic languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

multilingual embedding models
evaluation framework
evidence strength score
benchmark scarcity
cross-lingual transfer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Ana Gjorgjevikj
Computer Systems Department, Jožef Stefan Institute, Ljubljana, Slovenia
B
Barbara Koroušić Seljak
Computer Systems Department, Jožef Stefan Institute, Ljubljana, Slovenia
Tome Eftimov
Tome Eftimov
Computer Systems Department, Jožef Stefan Institute
StatisticsStochastic Optimization AlgorithmsMachine learningNatural Language Processing