🤖 AI Summary
To address the lack of systematic, large-scale evaluation benchmarks for low-resource languages like Turkish, this paper introduces TR-MMLU—the first comprehensive multitask language understanding benchmark for Turkish. It comprises 6,200 expert-validated multiple-choice questions spanning 62 subjects aligned with the Turkish national education curriculum. TR-MMLU enables fine-grained assessment of language comprehension, logical reasoning, and cross-domain knowledge acquisition. Systematic evaluation across leading open- and closed-weight large language models reveals substantial performance gaps in Turkish, particularly in domain-specific reasoning and long-range dependency modeling. TR-MMLU fills a critical gap in low-resource language model evaluation, providing a standardized, reproducible benchmark to guide future model development, data curation, and capability diagnostics. It establishes a new de facto standard for Turkish NLP evaluation.
📝 Abstract
Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations.