The Role of Prompt Language and Translation-Theory-Driven Prompts in Large Language Models: A Case Study on Spanish-Chinese Journalistic Translation

📅 2026-07-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of prompt design—integrating principles from fusion translation theory and varying prompt languages—on the quality of Spanish-to-Chinese news translation by large language models. Using GPT-5.2, the authors evaluate performance across 48 experimental conditions (comprising four prompt types, three prompt languages, and four editorial texts) through automatic metrics (BLEU and BERTScore-F1) and multidimensional human assessment via MQM. For the first time, translation theory is explicitly incorporated into prompt engineering, yielding significant improvements in expert human ratings, particularly in reducing “awkward style” errors. Results indicate that the BRIEF prompt achieves the highest MQM score (8.66 versus 7.84 for BASE), while the choice of prompt language exerts negligible influence, underscoring the critical role of theory-driven prompting in enhancing stylistic fluency in machine translation.
📝 Abstract
This study examines how prompt language and translation theory-driven prompt design influence the quality of Spanish-Chinese journalistic translations generated by GPT-5.2. A parallel corpus of four editorials from El Pais was translated under 48 experimental conditions (4 prompt types, 3 prompt languages, and 4 articles). Translation quality was assessed using BLEU and BERTScore-F1 for automated evaluation, alongside human evaluation based on the Multidimensional Quality Metrics (MQM) framework. Automated metrics identified the baseline prompt (BASE) as the best-performing condition, whereas human evaluation ranked the brief-oriented prompt (BRIEF) highest (MQM: 8.66 vs. 7.84), a reversal likely attributable to the single-reference constraint inherent in automated measures. Sub-error type analysis revealed that translation theory-driven prompts selectively reduced Awkward style errors, while Unidiomatic style errors persisted across conditions. Prompt language had a negligible impact under both evaluation paradigms. These results indicate that translation theory-driven prompts can yield measurable quality gains under expert evaluation of journalistic translations, although their pedagogical implications for language learners remain suggestive and require validation through user-based studies.
Problem

Research questions and friction points this paper is trying to address.

prompt language
translation theory
large language models
journalistic translation
translation quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

translation-theory-driven prompts
prompt design
Spanish-Chinese translation
MQM evaluation
large language models
H
Haohong Lai
Faculty of Translation and Interpreting, Autonomous University of Barcelona, Barcelona, Spain
W
Weijia Li
Faculty of Translation and Interpreting, Autonomous University of Barcelona, Barcelona, Spain