科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-14· cs.CL

ReMova: Fine-tuning LLMs for English to Belarusian translation

Mikita Pilinka, Aliaksandr Kliujeŭ, David Samuel, Yves Scherrer

原始摘要(英文原文)· Original abstract
This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet. A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system, supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ReMova: Fine-tuning LLMs for English to Belarusian translation — 科研速览 Science Skim