Institution profile

Zendesk

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?

Jun 08, 2026

This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.

0 citationsRead paper

Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages

Feb 02, 2026

This work addresses the confounding of genuine cross-lingual performance gaps with evaluation instability in multilingual model assessment. To isolate the intrinsic stability of evaluation methodologies, the authors propose a controlled generation paradigm that produces synthetic customer service dialogues in Estonian, Finnish, and Hungarian with identical parameters. Combining automatic metrics, LLM-as-a-judge evaluations, and native-speaker annotations, they find that surface-level measures—such as lexical diversity and semantic similarity—exhibit cross-lingual stability, whereas zero-shot judgments of pragmatic qualities like coherence and instruction following show substantial instability, including rank reversals and near-zero inter-language correlations. These findings indicate that current automatic evaluation methods require language-specific calibration for morphologically rich languages and underscore the value of the proposed paradigm as a diagnostic tool for cross-lingual evaluation robustness.

0 citationsRead paper
Recent publications

Latest Papers

How Far Can Prompting Go for Minimal-Edit Ukrainian Grammatical Error Correction?

Jun 08, 2026

This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.

0 citationsRead paper

Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages

Feb 02, 2026

This work addresses the confounding of genuine cross-lingual performance gaps with evaluation instability in multilingual model assessment. To isolate the intrinsic stability of evaluation methodologies, the authors propose a controlled generation paradigm that produces synthetic customer service dialogues in Estonian, Finnish, and Hungarian with identical parameters. Combining automatic metrics, LLM-as-a-judge evaluations, and native-speaker annotations, they find that surface-level measures—such as lexical diversity and semantic similarity—exhibit cross-lingual stability, whereas zero-shot judgments of pragmatic qualities like coherence and instruction following show substantial instability, including rank reversals and near-zero inter-language correlations. These findings indicate that current automatic evaluation methods require language-specific calibration for morphologically rich languages and underscore the value of the proposed paradigm as a diagnostic tool for cross-lingual evaluation robustness.

0 citationsRead paper