Trillion 7B Technical Report
This work addresses the challenge of developing efficient Korean-centric multilingual large language models (LLMs) under resource constraints. We propose Trillion-7B—the first trillion-parameter, Korean-hubbed multilingual LLM optimized for token efficiency. Methodologically, we introduce cross-lingual document attention (XLDA), integrated with language-aware data mixing, multilingual filtering, and a customized tokenizer, enabling efficient English knowledge transfer using only 2T training tokens—of which just 10% are multilingual (Korean, Japanese, Chinese). Experiments demonstrate state-of-the-art or highly competitive performance across 27 English, Korean, Japanese, and Chinese benchmarks, with significantly improved cross-lingual consistency. Full training requires only 59.4K H100 GPU-hours (≈$1.48M), achieving the highest token efficiency among existing Korean-centric multilingual LLMs.