Localizing AI: Evaluating Open-Weight Language Models for Languages of Baltic States

📅 2025-01-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of securely deploying large language models (LLMs) for low-resource Baltic languages—Lithuanian, Latvian, and Estonian—in privacy- and security-sensitive domains such as government and defense. We systematically evaluate open-source, locally deployable multilingual LLMs—including Llama 3, Gemma 2, Phi, and NeMo—across machine translation, multiple-choice question answering, and free-text generation. We identify pervasive token-level hallucinations (average error rate ≥5%, i.e., one error per 20 tokens), demonstrating that high translation accuracy does not guarantee semantic reliability. Through FP16/INT4 precision analysis and a custom evaluation benchmark, we find Gemma 2 approaches commercial-model performance, yet all models exhibit critical deficiencies requiring language-specific optimization. Our work establishes the first empirical benchmark for LLM deployment in privacy-critical, low-resource language settings and provides concrete, actionable pathways for improving linguistic fidelity and trustworthiness.

Technology Category

Application Category

📝 Abstract
Although large language models (LLMs) have transformed our expectations of modern language technologies, concerns over data privacy often restrict the use of commercially available LLMs hosted outside of EU jurisdictions. This limits their application in governmental, defence, and other data-sensitive sectors. In this work, we evaluate the extent to which locally deployable open-weight LLMs support lesser-spoken languages such as Lithuanian, Latvian, and Estonian. We examine various size and precision variants of the top-performing multilingual open-weight models, Llama~3, Gemma~2, Phi, and NeMo, on machine translation, multiple-choice question answering, and free-form text generation. The results indicate that while certain models like Gemma~2 perform close to the top commercially available models, many LLMs struggle with these languages. Most surprisingly, however, we find that these models, while showing close to state-of-the-art translation performance, are still prone to lexical hallucinations with errors in at least 1 in 20 words for all open-weight multilingual LLMs.
Problem

Research questions and friction points this paper is trying to address.

Privacy-Preserving
Low-Resource Languages
Language Model Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multilingual Models
Privacy Protection
Rare Language Support
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jurgita Kapovciute-Dzikiene
Tilde IT, Lithuania; Faculty of Informatics, Vytautas Magnus University, Lithuania
Toms Bergmanis
Toms Bergmanis
Tilde, University of Latvia
natural language processingmorphologymachine translationlanguage modeling
M
Marcis Pinnis
Tilde, Latvia; Faculty of Computing, University of Latvia