EuroAlpaca: Task-Preserving Localisation of Instruction Data for European Languages

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入EuroAlpaca方法,解决了机器翻译在扩展英语指令数据到多种欧洲语言时导致的任务关键约束和输出失真的问题。
📝 Abstract
Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and degrading models trained on such data. We introduce EuroAlpaca, a task-preserving localisation pipeline and near-parallel resource covering 50 European languages and regional varieties, together with European-IFEval, a multilingual benchmark for verifiable instruction following. Depending on the example, our pipeline applies field-wise MT while preserving task-critical content or reconstructs a task-equivalent target-language instance, followed by validation of cross-field coherence and target-language consistency. Across LoRA experiments with four LLMs, training on directly translated data improves ROUGE-L and F-BERT on the Aya Evaluation Suite, but reduces accuracy on European-IFEval by 29.8% relative to the unadapted baseline. In contrast, adaptation with EuroAlpaca improves accuracy by 12.9% over the same baseline, reversing the degradation caused by direct MT, while also achieving the highest ROUGE-L and F-BERT scores on Aya. These results show that preserving task semantics is essential for multilingual instruction tuning.
Problem

Research questions and friction points this paper is trying to address.

Machine Translation
Instruction-Tuning
Multilingual
Task-Preserving
Data Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-preserving localization
field-wise MT
cross-field coherence
European-IFEval
🔎 Similar Papers
No similar papers found.