Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

📅 2024-12-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Low-resource languages like Turkish lack systematic, comprehensive benchmarks for evaluating large language models (LLMs). Method: We introduce TR-MMLU—the first Turkish-language benchmark for LLM evaluation—comprising 6,200 high-quality, multiple-choice questions across 62 disciplines. Questions were curated via multi-source filtering, discipline-balanced sampling, expert human verification, and culturally grounded annotation; we also define a standardized evaluation protocol. Contribution/Results: TR-MMLU establishes the first reproducible, culturally adapted LLM evaluation framework for Turkish. Systematic evaluation of leading open and proprietary LLMs reveals critical impacts of tokenization strategies and fine-tuning approaches on performance, and identifies persistent cross-model bottlenecks in Turkish linguistic understanding. TR-MMLU provides a quantifiable, domain-diverse benchmark and actionable insights to guide model development, evaluation, and optimization for low-resource languages.

Technology Category

Application Category

📝 Abstract
Language models have made remarkable advancements in understanding and generating human language, achieving notable success across a wide array of applications. However, evaluating these models remains a significant challenge, particularly for resource-limited languages such as Turkish. To address this gap, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is constructed from a carefully curated dataset comprising 6200 multiple-choice questions across 62 sections, selected from a pool of 280000 questions spanning 67 disciplines and over 800 topics within the Turkish education system. This benchmark provides a transparent, reproducible, and culturally relevant tool for evaluating model performance. It serves as a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text and fostering the development of more robust and accurate language models. In this study, we evaluate state-of-the-art LLMs on TR-MMLU, providing insights into their strengths and limitations for Turkish-specific tasks. Our findings reveal critical challenges, such as the impact of tokenization and fine-tuning strategies, and highlight areas for improvement in model design. By setting a new standard for evaluating Turkish language models, TR-MMLU aims to inspire future innovations and support the advancement of Turkish NLP research.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Resource-scarce Languages
Turkish Language
Innovation

Methods, ideas, or system contributions that make the work stand out.

TR-MMLU
Turkish Multimodal Language Understanding
Large Language Model Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yıldız Technical University | Yeditepe University | Istanbul Bilgi University | Isik University
M
M. A. Bayram
Yıldız Technical University
Ali Arda Fincan
Ali Arda Fincan
Yeditepe University
A
Ahmet Semih Gumucs
Yeditepe University
B
Banu Diri
Yıldız Technical University
S
Savacs Yildirim
Istanbul Bilgi University
O
Oner Aytacs
Isik University