🤖 AI Summary
This study addresses the challenge of jointly modeling human evaluation criteria and cultural sensitivity—particularly Korean honorifics—in automatic literary translation assessment. We propose a two-stage fine-grained evaluation framework for English-to-Korean literary machine translation. Methodologically, we introduce the first explainable, fine-grained metrics integrating rule-based honorific and rhetorical detection with large language model (LLM)-driven semantic consistency scoring, enabling joint modeling of stylistic fidelity, politeness, and semantic accuracy. Our contributions are threefold: (1) significantly improved correlation with human judgments—outperforming BLEU, COMET, and other baselines; (2) empirical identification and quantification of systematic LLM preference biases toward specific translation variants; and (3) establishment of a critical performance ceiling—current automatic metrics still fall short of inter-annotator agreement—thereby providing an essential benchmark and concrete directions for future work.
📝 Abstract
In this work, we propose and evaluate the feasibility of a two-stage pipeline to evaluate literary machine translation, in a fine-grained manner, from English to Korean. The results show that our framework provides fine-grained, interpretable metrics suited for literary translation and obtains a higher correlation with human judgment than traditional machine translation metrics. Nonetheless, it still fails to match inter-human agreement, especially in metrics like Korean Honorifics. We also observe that LLMs tend to favor translations generated by other LLMs, and we highlight the necessity of developing more sophisticated evaluation methods to ensure accurate and culturally sensitive machine translation of literary works.