Larger Context Window, Fewer Overcorrections: Optimizing Prompts and Batching for Minimal-Edit Grammatical Error Correction
本文通过引入基于分类的指令、批量处理句子和优化提示的方法,解决了大型语言模型在最小编辑语法纠错任务中过度纠正的问题。
本文通过引入基于分类的指令、批量处理句子和优化提示的方法,解决了大型语言模型在最小编辑语法纠错任务中过度纠正的问题。
This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.
This work addresses the confounding of genuine cross-lingual performance gaps with evaluation instability in multilingual model assessment. To isolate the intrinsic stability of evaluation methodologies, the authors propose a controlled generation paradigm that produces synthetic customer service dialogues in Estonian, Finnish, and Hungarian with identical parameters. Combining automatic metrics, LLM-as-a-judge evaluations, and native-speaker annotations, they find that surface-level measures—such as lexical diversity and semantic similarity—exhibit cross-lingual stability, whereas zero-shot judgments of pragmatic qualities like coherence and instruction following show substantial instability, including rank reversals and near-zero inter-language correlations. These findings indicate that current automatic evaluation methods require language-specific calibration for morphologically rich languages and underscore the value of the proposed paradigm as a diagnostic tool for cross-lingual evaluation robustness.
本文通过引入基于分类的指令、批量处理句子和优化提示的方法,解决了大型语言模型在最小编辑语法纠错任务中过度纠正的问题。
This study investigates whether prompt engineering alone can approach the performance of fine-tuned large language models on Ukrainian minimal-edit grammatical error correction (GEC). We systematically evaluate twelve large language models—including eleven commercial systems and one open-source Ukrainian-specific model—using zero-shot and few-shot prompting, minimal-edit constraints, and model-assisted prompt optimization, enhanced with linguistically informed instructions grounded in Ukrainian grammar. Our work provides the first comprehensive validation of prompt engineering’s efficacy for Ukrainian GEC, revealing its strong dependence on prompt language and identifying five distinct overcorrection patterns tied to Ukrainian linguistic characteristics. The best-performing configuration, Gemini 1.5 Pro, achieves an F0.5 score of 69.22 on the UNLP 2023 benchmark, closing over 90% of the performance gap with the current fine-tuned state-of-the-art model.
This work addresses the confounding of genuine cross-lingual performance gaps with evaluation instability in multilingual model assessment. To isolate the intrinsic stability of evaluation methodologies, the authors propose a controlled generation paradigm that produces synthetic customer service dialogues in Estonian, Finnish, and Hungarian with identical parameters. Combining automatic metrics, LLM-as-a-judge evaluations, and native-speaker annotations, they find that surface-level measures—such as lexical diversity and semantic similarity—exhibit cross-lingual stability, whereas zero-shot judgments of pragmatic qualities like coherence and instruction following show substantial instability, including rank reversals and near-zero inter-language correlations. These findings indicate that current automatic evaluation methods require language-specific calibration for morphologically rich languages and underscore the value of the proposed paradigm as a diagnostic tool for cross-lingual evaluation robustness.