Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
提出SURE方法,基于奖励评估语法纠错效果,解决传统评估依赖参考答案和编辑重叠的问题,通过学习语法、忠实度和流畅性等标准来判断输出的有效性。
📝 Abstract
Grammatical error correction (GEC) evaluation has traditionally relied on reference or edit overlap, which can penalize valid rewrites that differ from gold corrections. Reference-free metrics reduce this dependence, but evaluating whether a fluent output is a valid correction of the source remains challenging. We propose SURE, a source-conditioned reward evaluator trained on within-source preferences spanning minimal-edit and rewrite-oriented corrections. SURE jointly learns an overall reward with criteria-level supervision for grammaticality, faithfulness, and fluency, together with span-level grounding for source-side error resolution. Experiments on SEEDA show that SURE performs competitively against strong baselines, with particular gains on rewrite-style corrections and more disentangled criteria-level diagnostics. Our code is available at https://github.com/hayeonggg/SURE.
Problem

Research questions and friction points this paper is trying to address.

Grammatical Error Correction
Reference-free Metrics
Rewrite-oriented Corrections
Innovation

Methods, ideas, or system contributions that make the work stand out.

SURE
reward-based evaluation
grammatical error correction
source-conditioned
H
Hayeong Ryu
Department of Artificial Intelligence, Chung-Ang University
S
Sunhee Jo
Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University
S
Seunguk Yu
Department of Artificial Intelligence, Chung-Ang University
YoungBin Kim
YoungBin Kim
Chung-Ang University
Machine Learning