Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient reliability of existing automatic evaluation metrics in historical language scenarios such as classical Chinese-to-English translation, which hinders their utility in digital humanities research. The authors propose the first diagnostic framework based on minimal pairs to systematically assess both reference-dependent and reference-free metrics with respect to their sensitivity to translation errors and tolerance for legitimate variants. Experimental results reveal that all evaluated metrics exhibit blind spots, though MetricX24 demonstrates superior performance overall. By uncovering the limitations of current metrics in cross-cultural historical translation contexts, this work introduces a diagnostic approach tailored to scholarly applications and lays the groundwork for developing more robust and interpretable evaluation tools.
📝 Abstract
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.
Problem

Research questions and friction points this paper is trying to address.

evaluation metrics
Classical Chinese translation
error detection
machine translation evaluation
digital humanities
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation metrics
Classical Chinese translation
minimal pairs
error sensitivity
reference-free evaluation
💼 Related Jobs
No related jobs found.
O
Osvaldo Quinjica
Department of Computer Science, University of Maryland, College Park
E
Eric Bennett
Department of Computer Science, University of Maryland, College Park; Department of East Asian Languages and Cultures, University of Maryland, College Park
Xinchen Yang
Xinchen Yang
PhD student of Computer Science, University of Maryland, College Park
Machine LearningTheoryLarge Language ModelsNatural Language Processing
A
Andrew Schonebaum
Department of East Asian Languages and Cultures, University of Maryland, College Park
Marine Carpuat
Marine Carpuat
Associate Professor, Computer Science, University of Maryland
Natural Language ProcessingMachine Translation