Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi

📅 2025-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tokenization challenge for morphologically rich, low-resource languages (e.g., Turkish). We propose the first multidimensional evaluation framework tailored to such languages, built upon the newly constructed TR-MMLU dataset. We systematically benchmark mainstream tokenizers across four dimensions: vocabulary size, token count, processing efficiency, and preservation of linguistic structure—introducing two novel metrics: language-specific tokenization ratio (%TR) and token purity (%Pure). Results show that %TR strongly correlates with downstream task performance, whereas %Pure exhibits weak correlation; scaling model parameters alone does not improve linguistic understanding. Crucially, this work provides the first empirical validation that %TR is a more effective indicator of tokenization quality than conventional metrics. It underscores the necessity of linguistically grounded, language-specific tokenization strategies for morphologically complex languages and establishes a reproducible evaluation paradigm and methodological foundation for low-resource NLP.

Technology Category

Application Category

📝 Abstract
Tokenization is a fundamental preprocessing step in Natural Language Processing (NLP), significantly impacting the capability of large language models (LLMs) to capture linguistic and semantic nuances. This study introduces a novel evaluation framework addressing tokenization challenges specific to morphologically-rich and low-resource languages such as Turkish. Utilizing the Turkish MMLU (TR-MMLU) dataset, comprising 6,200 multiple-choice questions from the Turkish education system, we assessed tokenizers based on vocabulary size, token count, processing time, language-specific token percentages (%TR), and token purity (%Pure). These newly proposed metrics measure how effectively tokenizers preserve linguistic structures. Our analysis reveals that language-specific token percentages exhibit a stronger correlation with downstream performance (e.g., MMLU scores) than token purity. Furthermore, increasing model parameters alone does not necessarily enhance linguistic performance, underscoring the importance of tailored, language-specific tokenization methods. The proposed framework establishes robust and practical tokenization standards for morphologically complex languages.
Problem

Research questions and friction points this paper is trying to address.

Evaluating tokenization impact on Turkish language models
Proposing metrics for linguistic structure preservation
Assessing correlation between tokenization and model performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel evaluation framework for Turkish tokenization
Language-specific metrics like %TR and %Pure
Tailored tokenization for morphologically-rich languages
🔎 Similar Papers
No similar papers found.
M. Ali Bayram
M. Ali Bayram
Yıldız Teknik Üniversitesi
LLMAI
Ali Arda Fincan
Ali Arda Fincan
Yeditepe University
Ahmet Semih Gümüş
Ahmet Semih Gümüş
Yeditepe University
S
Sercan Karakaş
The University of Chicago, Chicago, IL, USA
B
Banu Diri
Yıldız Technical University, Istanbul, Turkey
S
Savaş Yıldırım
Istanbul Bilgi University, Istanbul, Turkey