Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

📅 2025-02-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit poor tokenization quality and inadequate preservation of morphological structure for typologically rich, low-resource languages—e.g., Turkish—due to suboptimal subword segmentation. Method: We propose the first multidimensional tokenization evaluation framework tailored to such languages, built upon a linguistically adapted Turkish version of MMLU (6,200 items). Our framework systematically assesses five dimensions: vocabulary size, token count, inference latency, percentage of target-language valid tokens (%TR), and token purity. Crucially, we formally define and empirically validate %TR as a more predictive proxy for downstream performance than token purity. Results: Experiments reveal that scaling model parameters does not inherently improve tokenization quality; instead, linguistic alignment should supersede scale-driven design. We establish %TR as the core evaluation metric and deliver a reproducible, extensible tokenization paradigm for low-resource languages—thereby advancing LLM preprocessing methodology.

Technology Category

Application Category

📝 Abstract
Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. This paper introduces a novel framework for systematically evaluating tokenization strategies, addressing challenges in morphologically rich and low-resource languages. Using a Turkish dataset of 6,200 multiple-choice questions from the Massive Multitask Language Understanding (MMLU) benchmark, the framework assesses tokenizers across five key metrics: vocabulary size, token count, processing time, language-specific token percentages (%TR), and token purity. These metrics provide a structured approach to evaluating how well tokenizers preserve linguistic structures. While %TR measures the proportion of valid words in the target language, %Pure assesses the alignment of tokens with meaningful linguistic units, such as roots and valid morphemes, minimizing semantic fragmentation. The findings reveal that %TR, introduced as a critical metric, exhibits a stronger correlation with downstream performance (e.g., MMLU scores) than token purity, emphasizing its role in improving model accuracy. Additionally, larger model parameters do not necessarily yield better tokenization quality or enhanced results, highlighting the importance of tailored tokenization strategies that prioritize linguistic alignment. This framework sets a new standard for developing robust tokenization methods optimized for morphologically complex and low-resource languages. Future work will refine morphological analysis, explore domain-specific customizations, and conduct cross-linguistic evaluations to further enhance tokenization practices.
Problem

Research questions and friction points this paper is trying to address.

Evaluates tokenization in morphologically rich languages
Introduces metrics to assess linguistic structure preservation
Focuses on Turkish for low-resource language tokenization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novel framework for tokenization evaluation
Turkish dataset with 6,200 questions
Five metrics for tokenizer assessment
🔎 Similar Papers
2024-06-21arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
M
M. A. Bayram
Yıldız Technical University
Ali Arda Fincan
Ali Arda Fincan
Yeditepe University
A
Ahmet Semih Gumucs
Yeditepe University
S
Sercan Karakacs
University of Chicago
B
Banu Diri
Yıldız Technical University
S
Savacs Yildirim
Istanbul Bilgi University