Institution profile

Institute of the Estonian Language

Academic institutioneurope · ee
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

Mar 02, 2026

This work addresses the significant performance gap of multilingual large language models on low-resource languages such as Estonian, while maintaining strong capabilities in high-resource languages and general tasks. Building upon Llama 3.1 8B, the authors propose a balanced multilingual data mixing strategy for continued pretraining, augmented with English replay and enriched with code, mathematical, and instructional data. The model is further aligned through supervised fine-tuning, preference optimization, and chat vector fusion techniques. This approach yields substantial improvements across Estonian language understanding, knowledge recall, reasoning, translation, and instruction-following benchmarks, while preserving competitive performance on English and general-purpose evaluations, thereby achieving an effective balance between low-resource language enhancement and overall multilingual competence.

0 citationsRead paper

Vision-Enabled LLMs in Historical Lexicography: Digitising and Enriching Estonian-German Dictionaries from the 17th and 18th Centuries

Oct 09, 2025

This study addresses the challenges of digitizing and semantically enriching 17th–18th-century Estonian Gothic-type (Fraktur) printed dictionaries. Methodologically, it proposes an end-to-end solution based on multimodal large language models (MLLMs), featuring a dual-model collaborative framework for overlapping image patch processing. The approach integrates visual understanding, context-aware prompting, and JSON-structured output generation to enable zero-shot paleographic text recognition and lexical entry structuring. Additionally, a cross-source unified data model is introduced to support modern semantic mapping and lexical gap filling for historical vocabulary. Experimental results show a 41% error-free structuring rate for headwords and an 81% accuracy in semantic completion—marking substantial improvements in automation efficiency and scalability for low-resource historical linguistic corpora. The work establishes a novel paradigm for intelligent curation of endangered language heritage.

0 citationsRead paper
Recent publications

Latest Papers

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

Mar 02, 2026

This work addresses the significant performance gap of multilingual large language models on low-resource languages such as Estonian, while maintaining strong capabilities in high-resource languages and general tasks. Building upon Llama 3.1 8B, the authors propose a balanced multilingual data mixing strategy for continued pretraining, augmented with English replay and enriched with code, mathematical, and instructional data. The model is further aligned through supervised fine-tuning, preference optimization, and chat vector fusion techniques. This approach yields substantial improvements across Estonian language understanding, knowledge recall, reasoning, translation, and instruction-following benchmarks, while preserving competitive performance on English and general-purpose evaluations, thereby achieving an effective balance between low-resource language enhancement and overall multilingual competence.

0 citationsRead paper

Vision-Enabled LLMs in Historical Lexicography: Digitising and Enriching Estonian-German Dictionaries from the 17th and 18th Centuries

Oct 09, 2025

This study addresses the challenges of digitizing and semantically enriching 17th–18th-century Estonian Gothic-type (Fraktur) printed dictionaries. Methodologically, it proposes an end-to-end solution based on multimodal large language models (MLLMs), featuring a dual-model collaborative framework for overlapping image patch processing. The approach integrates visual understanding, context-aware prompting, and JSON-structured output generation to enable zero-shot paleographic text recognition and lexical entry structuring. Additionally, a cross-source unified data model is introduced to support modern semantic mapping and lexical gap filling for historical vocabulary. Experimental results show a 41% error-free structuring rate for headwords and an 81% accuracy in semantic completion—marking substantial improvements in automation efficiency and scalability for low-resource historical linguistic corpora. The work establishes a novel paradigm for intelligent curation of endangered language heritage.

0 citationsRead paper