EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
This work addresses the significant performance gap of multilingual large language models on low-resource languages such as Estonian, while maintaining strong capabilities in high-resource languages and general tasks. Building upon Llama 3.1 8B, the authors propose a balanced multilingual data mixing strategy for continued pretraining, augmented with English replay and enriched with code, mathematical, and instructional data. The model is further aligned through supervised fine-tuning, preference optimization, and chat vector fusion techniques. This approach yields substantial improvements across Estonian language understanding, knowledge recall, reasoning, translation, and instruction-following benchmarks, while preserving competitive performance on English and general-purpose evaluations, thereby achieving an effective balance between low-resource language enhancement and overall multilingual competence.