🤖 AI Summary
This study addresses the insufficient reliability of existing automatic evaluation metrics in historical language scenarios such as classical Chinese-to-English translation, which hinders their utility in digital humanities research. The authors propose the first diagnostic framework based on minimal pairs to systematically assess both reference-dependent and reference-free metrics with respect to their sensitivity to translation errors and tolerance for legitimate variants. Experimental results reveal that all evaluated metrics exhibit blind spots, though MetricX24 demonstrates superior performance overall. By uncovering the limitations of current metrics in cross-cultural historical translation contexts, this work introduces a diagnostic approach tailored to scholarly applications and lays the groundwork for developing more robust and interpretable evaluation tools.
📝 Abstract
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.