Mi:dm 2.0 Korea-centric Bilingual Language Models
This work addresses the limitations of current large language models in handling Korean, which stem from low-quality training data and a lack of cultural alignment, hindering their ability to capture Korea-specific values, commonsense knowledge, and nuanced emotional expressions. To overcome these challenges, we propose Mi:dm 2.0—the first bilingual large language model systematically integrating Korean sociocultural commonsense and reasoning patterns. Through high-quality data curation, synthetic data generation, a curriculum learning–guided data mixing strategy, and a Korean-optimized tokenizer, Mi:dm 2.0 achieves deep contextual understanding of local nuances. Released under the MIT License in both general and lightweight variants, the model attains state-of-the-art zero-shot performance on Korean benchmarks such as KMMLU, significantly outperforming existing models and advancing the development of the K-intelligence ecosystem.