MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文开发了爱尔兰语的分词资源MoirfEolas及评价指标CríochScore,评估了不同分词算法与形态边界的一致性,发现Unigram语言模型表现最佳。
📝 Abstract
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective eclipses, prefixes and suffixes as well as an evaluation metric CríochScore, that evaluates the alignment of tokenizations with the morphological boundaries present in MoirfEolas. We evaluate common tokenization algorithms using CríochScore as well as intrinsic metrics present in the tokenization literature. We find that the Unigram Language Model aligns with Irish morphology more often than the other algorithms evaluated. We also find trade-offs between morphological-alignment of tokenization with both compression as well as vocabulary efficiency, providing practical insights for Irish natural language processing development. This dataset contributes towards combating the Irish language's low-resource status; moreover, the construction process reported in this paper can be emulated by other languages to create specialised morphological resources.
Problem

Research questions and friction points this paper is trying to address.

Tokenization
Morphology
Irish Language
Natural Language Processing
Resource Development
Innovation

Methods, ideas, or system contributions that make the work stand out.

MoirfEolas
CríochScore
Tokenization Alignment
Irish Morphology
Unigram Language Model
🔎 Similar Papers
2024-06-21arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
J
Jane Adkins
ADAPT Centre, Dublin City University
A
Abigail Walsh
ADAPT Centre, Dublin City University
B
Brian Davis
ADAPT Centre, Dublin City University
Elaine Uí Dhonnchadha
Elaine Uí Dhonnchadha
Trinity College Dublin
computational linguisticsnatural language processingcorpus linguisticsIrishCALL