Last Translation Benchmark

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有翻译模型评估方法的局限性,提出Last Translation Benchmark,通过人工设计并同行评审的例子及手写验证规则来更可靠地评估模型。
📝 Abstract
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.
Problem

Research questions and friction points this paper is trying to address.

machine translation
benchmark
evaluation methods
reliability
actionable assessments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Last Translation Benchmark
handcrafted verification rules
reliable and actionable evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Vilém Zouhar
Vilém Zouhar
PhD, ETH Zürich
Natural Language ProcessingQuality EstimationMachine Translation
Niyati Bafna
Niyati Bafna
Johns Hopkins University, Center for Language and Speech Processing
Low-resource NLPLarge Language ModellingMachine TranslationBilingual Lexicon Induction
Mukund Choudhary
Mukund Choudhary
Mohamed Bin Zayed University of Artificial Intelligence
Computational LinguisticsMetalinguisticsLLMsCognitive ScienceNatural Language Processing
Maike Züfle
Maike Züfle
Karlsruhe Institute of Technology - KIT
Sara Rajaee
Sara Rajaee
Ph.D. student at University of Amsterdam
Natural language processingArtificial Intelligence
Pinzhen Chen
Pinzhen Chen
University of Edinburgh
large language modelsLLM post-trainingmachine translationmultilinguality
Jannis Vamvas
Jannis Vamvas
University of Zurich
Sara Papi
Sara Papi
Researcher at FBK
Speech ProcessingSpeech TranslationMultimodal LLM
Ona de Gibert
Ona de Gibert
PhD Student @ University of Helsinki
Machine TranslationMultilingualityKnowledge Distillation
Bhavitvya Malik
Bhavitvya Malik
Research Assistant, University of Edinburgh
natural language processingspeech
Eliya Habba
Eliya Habba
Hebrew University of Jerusalem
Orfeas Menis Mastromichalakis
Orfeas Menis Mastromichalakis
PhD Student, National Technical University of Athens
Explainable AIAI EthicsNLP
Patrícia Schmidtová
Patrícia Schmidtová
Institute of Formal and Applied Linguistics, Charles University
Natural Language Processing
M
Michelle Wastl
Sheriff Issaka
Sheriff Issaka
UCLA
Natural Language ProcessingMultilingual NLPLow-resource NLP
Leshem Choshen
Leshem Choshen
MIT, IBM AI research
Model RecyclingEvolving Collaborative PretrainingEvaluationModel MergingOpen the Black Box
Stella Biderman
Stella Biderman
EleutherAI
Natural Language ProcessingArtificial IntelligenceLanguage ModelingDeep Learning
Antonis Anastasopoulos
Antonis Anastasopoulos
Assistant Professor, George Mason University
Natural Language ProcessingMachine TranslationSpeech Recognition
Jan Niehues
Jan Niehues
Institute for Anthropomatics and Robotics (IAR), Karlsruhe Institute for Technology (KIT)
natural language processing - machine translation
R
Rico Sennrich
Mrinmaya Sachan
Mrinmaya Sachan
Assistant Professor, ETH Zürich
Natural Language ProcessingReasoningAI for Education
Ondřej Bojar
Ondřej Bojar
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics
machine translationspeech translationparsingtreebanking
Kenton Murray
Kenton Murray
Research Scientist, Johns Hopkins
Machine LearningNatural Language ProcessingMachine TranslationSemanticsNeural Networks
Jörg Tiedemann
Jörg Tiedemann
Professor of Language Technology, University of Helsinki
computational linguisticsmachine translationmachine learningnatural language processinginformation retrieval
Alham Fikri Aji
Alham Fikri Aji
MBZUAI, Monash Indonesia
MultilingualityLow-resource NLPLanguage ModelingMachine Translation