SwiLTra-Bench: The Swiss Legal Translation Benchmark

📅 2025-03-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Swiss multilingual legal translation has long relied on scarce bilingual legal linguists, impeding judicial accessibility. To address this, we introduce SwiLTra-Bench—the first large-scale, expert-validated benchmark for legal translation across five official Swiss languages (German, French, Italian, Romansh, and English), comprising 180,000 aligned sentence pairs from statutes, case summaries, and press releases. We propose a standardized evaluation framework tailored to Switzerland’s federal multilingual legal system and develop SwiLTra-Judge, an automated assessment system achieving a 0.92 Spearman correlation with human expert judgments. Experimental results show that state-of-the-art closed-source LLMs (e.g., Claude-3.5-Sonnet) achieve superior zero-shot translation performance across all language directions; while open-source models improve substantially after fine-tuning, they still underperform the best zero-shot baselines. This work establishes a rigorous empirical foundation—comprising a high-quality benchmark, a domain-specific evaluation protocol, and validated automatic metrics—for trustworthy legal translation in multilingual jurisdictions.

Technology Category

Application Category

📝 Abstract
In Switzerland legal translation is uniquely important due to the country's four official languages and requirements for multilingual legal documentation. However, this process traditionally relies on professionals who must be both legal experts and skilled translators -- creating bottlenecks and impacting effective access to justice. To address this challenge, we introduce SwiLTra-Bench, a comprehensive multilingual benchmark of over 180K aligned Swiss legal translation pairs comprising laws, headnotes, and press releases across all Swiss languages along with English, designed to evaluate LLM-based translation systems. Our systematic evaluation reveals that frontier models achieve superior translation performance across all document types, while specialized translation systems excel specifically in laws but under-perform in headnotes. Through rigorous testing and human expert validation, we demonstrate that while fine-tuning open SLMs significantly improves their translation quality, they still lag behind the best zero-shot prompted frontier models such as Claude-3.5-Sonnet. Additionally, we present SwiLTra-Judge, a specialized LLM evaluation system that aligns best with human expert assessments.
Problem

Research questions and friction points this paper is trying to address.

Addresses bottlenecks in Swiss legal translation due to multilingual requirements.
Evaluates LLM-based systems for translating Swiss legal documents across languages.
Compares performance of specialized and frontier models in legal translation tasks.
Innovation

Methods, ideas, or system contributions that make the work stand out.

SwiLTra-Bench: multilingual legal translation benchmark
Fine-tuning SLMs improves translation quality significantly
SwiLTra-Judge: LLM evaluation system for expert alignment
🔎 Similar Papers
No similar papers found.
Joel Niklaus
Joel Niklaus
Hugging Face, Stanford University
Natural Language ProcessingLegal NLPLegal AI
J
Jakob Merane
ETH Zurich, Max Planck Institute for Research on Collective Goods
Luka Nenadic
Luka Nenadic
Ph.D. Student, ETH Zurich (Center for Law & Economics)
Law and Economics
Sina Ahmadi
Sina Ahmadi
University of Zurich
Natural Language ProcessingComputational Linguistics
Y
Yingqiang Gao
University of Zurich
C
Cyrill A. H. Chevalley
University of Basel
C
Claude Humbel
University of Zurich
C
Christophe Gosken
ETH Zurich
L
Lorenzo Tanzi
University of Geneva
T
Thomas Luthi
Canton of Solothurn
S
Stefan Palombo
Harvey
S
Spencer Poff
Harvey
Boling Yang
Boling Yang
University of Washington
Artificial IntelligenceReinforcement LearningRobotics
N
Nan Wu
Harvey
M
Matthew Guillod
Harvey
R
Robin Mami'e
Swiss Federal Supreme Court
Daniel Brunner
Daniel Brunner
CNRS researcher, FEMTO-ST, Optics department, Besancon
Photonic neural networksunconventional computationsemiconductor nonlinear opticscomplex photonicsnonlinear dynamics
J
Julio Pereyra
Harvey
N
Niko A. Grupen
Harvey