Dicta-LM 3.0: Advancing The Frontier of Hebrew Sovereign LLMs
This work addresses the scarcity of high-quality open-source large language models (LLMs) for low-resource languages like Hebrew, which hinders localized applications. We present the first systematic effort to develop a family of Hebrew LLMs at three scales—1.7B, 12B, and 24B parameters—based on Mistral-Small-3.1, NVIDIA Nemotron Nano V2, and Qwen3-1.7B, respectively. These models are adapted using large-scale Hebrew–English mixed corpora, support 65K-token context lengths and tool calling, and are released in both base and chat variants. Additionally, we introduce the first comprehensive evaluation benchmark for Hebrew chat models, demonstrating strong performance across tasks including translation, summarization, Winograd schema resolution, Israeli commonsense question answering, and nikud (vowel diacritic) restoration. The proposed framework is readily generalizable to other non-English languages and significantly advances Hebrew natural language processing.