Representing and Parsing Korean Constituency Structure at Different Levels of Granularity

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了不同粒度下韩语成分结构的表示与解析问题,通过比较三种基于Penn Korean Treebank的解析表示方法,并使用非二进制转换型成分解析器进行评估。
📝 Abstract
Korean constituency parsing raises a representational challenge because the terminal units of a phrase-structure tree do not straightforwardly correspond to simple surface words. Korean eojeols are morphologically complex spacing units, and existing constituency resources differ in how they represent eojeol-internal morphology and non-overt elements. This paper compares three constituency parsing representations derived from the Penn Korean Treebank: Morpheme+XPOS, Eojeol+XPOS, and Eojeol+UPOS. We construct these representations by removing null elements, aligning Penn Korean phrase structure with overt eojeol tokens, preserving Penn Korean phrase labels where possible, and varying the terminal and preterminal layers. We then evaluate canonical non-binary transition-based constituency parsers in top-down, in-order, and bottom-up orders under a shared modeling and evaluation setup. All experiments use gold terminal segmentation and gold preterminal labels and therefore evaluate constituency parsing conditioned on gold morphosyntactic annotation. Eojeol terminals yield shorter transition sequences, but Eojeol+UPOS parsing substantially underperforms the morphologically richer conditions. Eojeol+XPOS narrows this gap, while Morpheme+XPOS gives the strongest results even after its predictions are projected to the eojeol terminal domain. Under these gold-annotation conditions, the results show that fine-grained morphological and XPOS representations provide valuable evidence for the evaluated parsers. This empirical finding concerns the information available for parsing and does not by itself determine the linguistically preferable terminal domain. Independently, linguistic and resource-design considerations motivate eojeol as a stable and interpretable surface domain for phrase-structure annotation, with morpheme-level and XPOS information retained as aligned morphosyntactic evidence.
Problem

Research questions and friction points this paper is trying to address.

Korean constituency parsing
representational challenge
eojeols
morphological complexity
phrase-structure tree
Innovation

Methods, ideas, or system contributions that make the work stand out.

Morpheme+XPOS
Eojeol+XPOS
Eojeol+UPOS
constituency parsing
morphological representation
🔎 Similar Papers
No similar papers found.
J
Jungyeul Park
Korea Advanced Institute of Science & Technology, South Korea
KyungTae Lim
KyungTae Lim
École normale supérieure
Natural Language Processing
Z
Zihao Huang
The University of British Columbia, Canada
E
Eunkyul Leah Jo
The University of British Columbia, Canada
Yige Chen
Yige Chen
College of Computer Science and Artificial Intelligence, Wenzhou University
Networking
C
Chulwoo Park
Anyang University, South Korea