ReMova: Fine-tuning LLMs for English to Belarusian translation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对英语到白俄罗斯语的机器翻译,通过特定的数据清洗流程和微调方法解决了白俄罗斯语正字法不一致、数据噪声等问题,提高了翻译质量。
📝 Abstract
This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet. A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system, supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.
Problem

Research questions and friction points this paper is trying to address.

English-Belarusian translation
data quality
orthographies
noise
misspelling
Innovation

Methods, ideas, or system contributions that make the work stand out.

data-cleaning pipeline
orthographies
noise in training data