SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
This study addresses the challenge of evaluating subword tokenizers for morphologically rich languages like those in the Uralic family, where high-quality morpheme lexicons are scarce. To circumvent reliance on corpora, the authors propose a method grounded in the Minimum Description Length (MDL) principle, incorporating a self-referential atomicity scoring mechanism that leverages intraword structural cues to filter out compounds and construct high-purity morpheme lexicons. The work delivers the first empirically supported morpheme resources for Finnish, Hungarian, and Estonian, introduces an Integrated Performance Score (IPS) to balance coverage against over-segmentation, and employs elbow-point analysis to recommend optimal BPE vocabulary sizes. Findings highlight the limitations of standard BPE in highly agglutinative languages. The accompanying SampoNLP toolkit and resources are publicly released.