AI agents may be worth the hype but not the resources (yet): An initial exploration of machine translation quality and costs in three language pairs in the legal and news domains

📅 2025-05-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study empirically evaluates the quality, cost, and efficiency of LLMs and multi-agent coordination for English-to-Spanish/Catalan/Turkish machine translation in legal and news domains, benchmarking against Google Translate, GPT-4o, o1-preview, and two GPT-4o–driven agent workflows. Methodologically, it integrates automatic metrics (COMET, BLEU, chrF2, TER), expert human evaluation (adequacy and fluency), and token-level cost modeling based on April 2025 pricing. The work reveals, for the first time, a significant misalignment between surface-level automatic scores and human judgments. It proposes a cost-aware, multidimensional evaluation framework and identifies lightweight coordination, selective activation, and hybrid pipelines as key optimization levers. Results show: (1) Neural MT dominates most automatic metrics; (2) o1-preview achieves best performance across 5/6 human evaluation dimensions; (3) iterative agent systems occasionally surpass o1-preview but incur up to 15× higher token costs than NMT—highlighting a critical cost-efficiency bottleneck in current agent-based translation.

Technology Category

Application Category

📝 Abstract
Large language models (LLMs) and multi-agent orchestration are touted as the next leap in machine translation (MT), but their benefits relative to conventional neural MT (NMT) remain unclear. This paper offers an empirical reality check. We benchmark five paradigms, Google Translate (strong NMT baseline), GPT-4o (general-purpose LLM), o1-preview (reasoning-enhanced LLM), and two GPT-4o-powered agentic workflows (sequential three-stage and iterative refinement), on test data drawn from a legal contract and news prose in three English-source pairs: Spanish, Catalan and Turkish. Automatic evaluation is performed with COMET, BLEU, chrF2 and TER; human evaluation is conducted with expert ratings of adequacy and fluency; efficiency with total input-plus-output token counts mapped to April 2025 pricing. Automatic scores still favour the mature NMT system, which ranks first in seven of twelve metric-language combinations; o1-preview ties or places second in most remaining cases, while both multi-agent workflows trail. Human evaluation reverses part of this narrative: o1-preview produces the most adequate and fluent output in five of six comparisons, and the iterative agent edges ahead once, indicating that reasoning layers capture semantic nuance undervalued by surface metrics. Yet these qualitative gains carry steep costs. The sequential agent consumes roughly five times, and the iterative agent fifteen times, the tokens used by NMT or single-pass LLMs. We advocate multidimensional, cost-aware evaluation protocols and highlight research directions that could tip the balance: leaner coordination strategies, selective agent activation, and hybrid pipelines combining single-pass LLMs with targeted agent intervention.
Problem

Research questions and friction points this paper is trying to address.

Comparing machine translation quality of LLMs vs NMT systems
Evaluating cost-efficiency of multi-agent workflows in translation
Assessing human vs automatic metrics for translation performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarking five machine translation paradigms
Evaluating with automatic and human metrics
Proposing cost-aware hybrid pipelines
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.