Targeted Syntactic Evaluation of Language Models on Georgian Case Alignment
This study evaluates the ability of language models to handle nominative, ergative, and dative case marking within the split-ergative system of Georgian—a low-resource language with semantically complex ergative alignment. Leveraging a treebank and the Grew query language, the authors construct a fine-grained evaluation set comprising 370 minimal pairs across seven syntactic tasks (50–70 instances each). They systematically test five encoder and two decoder models, establishing the first syntactic benchmark for Georgian and proposing a methodology generalizable to other low-resource languages. The high-quality test set is publicly released. Results reveal that model performance strongly correlates with case form frequency (NOM > DAT > ERG), with ergative marking consistently the weakest, highlighting the joint impact of data scarcity and the semantic intricacies of ergativity on model accuracy.