Score
Designing and applying quantitative and qualitative evaluations to compare machine or human translations on specialised texts, and operationalizing literary/critical frameworks (e.g., macrostructure, circulation, untranslatability) into measurable criteria for system comparison.
This study addresses the longstanding reliance on subjective judgment in narrative quality assessment by introducing a computational framework grounded in 33 quantifiable linguistic features spanning lexical, syntactic, and semantic dimensions. For the first time, this work systematically applies multidimensional quantitative stylometric indicators to the automatic evaluation of narrative quality. Leveraging natural language processing, clustering analysis, and similarity matrix construction, the proposed model achieves near-perfect discrimination between texts authored by professional editors and self-published writers. Furthermore, it significantly outperforms existing evaluation metrics on a manually annotated dataset, thereby overcoming the limitations inherent in traditional story-level assessment approaches.
Existing automatic evaluation metrics and large language models employed as judges (LLM-as-a-judge) struggle to accurately assess creativity in literary translation and exhibit significant discrepancies with professional human evaluations. This study constructs a literary translation dataset spanning three modalities, genres, and language pairs, annotated by professional translators for fine-grained creative features, and systematically evaluates the performance of both traditional automatic metrics and LLMs along this dimension. The work reveals, for the first time, that LLMs display a systematic preference for machine-translated outputs and frequently misjudge culturally appropriate creative expressions—particularly in highly literary texts such as poetry, where correlations between current evaluation methods and human judgments markedly decline. These findings underscore the urgent need for novel evaluation frameworks capable of accommodating unconventional yet valid literary translations.
This study addresses the challenge of evaluating literary machine translation (MT) quality. We introduce LITEVAL-CORPUS—the first paragraph-level, human-validated parallel corpus for literary translation—covering four language pairs, 2,000+ translations, and 13,000+ sentences, enabling systematic comparison of human and large language model (LLM) output. Our key contributions include: (1) empirical evidence that the widely adopted Multidimensional Quality Metrics (MQM) framework exhibits severe miscalibration in literary contexts, whereas Best–Worst Scaling (BWS), a simple preference-ranking method, achieves 80–100% accuracy in identifying high-quality human translations; (2) student annotators applying MQM incur ~60% misclassification rates; (3) all automatic metrics attain ≤20% accuracy; and (4) human translations consistently surpass state-of-the-art LLMs in literariness and stylistic diversity, with LLMs exhibiting over-reliance on literal rendering and limited creative adaptation. The work establishes a robust, efficient human evaluation paradigm tailored to literary MT and provides a critical benchmark for future automatic evaluation research.
This study addresses the challenge of jointly modeling human evaluation criteria and cultural sensitivity—particularly Korean honorifics—in automatic literary translation assessment. We propose a two-stage fine-grained evaluation framework for English-to-Korean literary machine translation. Methodologically, we introduce the first explainable, fine-grained metrics integrating rule-based honorific and rhetorical detection with large language model (LLM)-driven semantic consistency scoring, enabling joint modeling of stylistic fidelity, politeness, and semantic accuracy. Our contributions are threefold: (1) significantly improved correlation with human judgments—outperforming BLEU, COMET, and other baselines; (2) empirical identification and quantification of systematic LLM preference biases toward specific translation variants; and (3) establishment of a critical performance ceiling—current automatic metrics still fall short of inter-annotator agreement—thereby providing an essential benchmark and concrete directions for future work.
Professional-domain machine translation often neglects communicative goals and client requirements—i.e., translation norms—leading to outputs misaligned with real-world practice. Method: This study pioneers the systematic integration of translation norm theory into machine translation, proposing a norm-explicit LLM-based translation framework. It leverages prompt engineering to elicit multi-style translations from large language models and employs a multidimensional evaluation combining expert error analysis, user preference ranking, and automated metrics. Results: Evaluated on investor relations texts from 33 listed companies, the method consistently outperforms official human translations in human evaluation. Its core contributions are: (1) establishing a norm-driven translation paradigm; (2) empirically validating that norm-guided MT can surpass professional human translation; and (3) providing an interpretable, controllable pathway for business-oriented machine translation.
This study investigates whether machine translation preserves the complexity of source texts and examines the relationship between textual complexity and translation difficulty. Building upon the Common European Framework of Reference for Languages (CEFR), the authors propose the first evaluation framework to analyze the interaction between text complexity and machine translation performance. They assess a range of open-source, closed-source, and commercial translation systems across six languages in terms of their ability to maintain or shift CEFR-based complexity levels. The findings reveal that texts at higher CEFR levels are more challenging to translate accurately, and that translated outputs frequently exhibit significant deviations from the original CEFR complexity ratings across most languages. This work provides novel empirical insights and quantitative methods for estimating translation difficulty and generating multilingual educational content.
This study investigates the development of students’ critical evaluation capabilities in AI-assisted translation pedagogy by engaging them in comparative assessments of outputs from general-purpose large language models and online machine translation systems. Participants performed post-editing tasks and justified their decisions using both automatic metrics (e.g., BLEU) and human evaluations focusing on fluency and accuracy. Findings reveal that students primarily based their judgments on multidimensional criteria—including terminological precision, linguistic naturalness, and anticipated editing effort—and often preferred translations that diverged from rankings suggested by automatic scores, demonstrating a nuanced, critical judgment that transcends quantitative metrics. This work challenges conventional paradigms of system evaluation by uncovering the complexity and pedagogical significance of learners’ subjective assessment logic in authentic instructional contexts.
Current automatic Machine Translation Quality Estimation (QE) systems lack reliability in real-world scenarios because their segment-level evaluations neglect critical dimensions such as discourse coherence, stylistic consistency, and rhetorical adequacy. Through theoretical analysis and empirical investigation, this study systematically uncovers structural limitations in QE—particularly concerning generalization capacity, data bias, overfitting, annotation noise, and the modeling of linguistic complexity—and, for the first time, identifies an inherent bottleneck rooted in the cognitive foundations of translation itself. These findings challenge the prevailing assumption that QE performance can be sufficiently improved merely by scaling up data or model capacity. The paper cautions against deploying existing QE systems as the sole basis for decision routing or bypassing human review in production environments and advocates redirecting future research toward automating human evaluation grounded in the Multidimensional Quality Metrics (MQM) framework.
This study investigates the quality differences between machine translation (MT) and human post-editing (PE) in domain-specific English-to-French translation tasks, as well as the influence of post-editors’ professional backgrounds on editing outcomes. The experiment compares three leading MT systems—DeepL, eTranslation, and Systran—and involves two groups of post-editors: linguists/translators and NLP experts. A fine-grained error taxonomy tailored to MT and PE is employed for manual evaluation. For the first time in a specialized translation setting, the interaction between MT system performance and editor background is jointly examined. Findings reveal that terminological accuracy and linguistic fluency are significantly affected by domain expertise, highlighting current limitations of MT in handling specialized language and underscoring the value of interdisciplinary post-editing teams.
This study investigates the impact of prompt design—integrating principles from fusion translation theory and varying prompt languages—on the quality of Spanish-to-Chinese news translation by large language models. Using GPT-5.2, the authors evaluate performance across 48 experimental conditions (comprising four prompt types, three prompt languages, and four editorial texts) through automatic metrics (BLEU and BERTScore-F1) and multidimensional human assessment via MQM. For the first time, translation theory is explicitly incorporated into prompt engineering, yielding significant improvements in expert human ratings, particularly in reducing “awkward style” errors. Results indicate that the BRIEF prompt achieves the highest MQM score (8.66 versus 7.84 for BASE), while the choice of prompt language exerts negligible influence, underscoring the critical role of theory-driven prompting in enhancing stylistic fluency in machine translation.