科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ MonTi Monografías de Traducción e Interpretación2026-05-27· Computer science

ChatGPT vs. DeepL vs. Google Translate: a human evaluation of multiword expressions’ machine translation quality

Carlos Manuel Hidalgo Ternero, Vicent Briva-Iglesias

原始摘要(英文原文)· Original abstract
Multiword expressions (MWEs) remain a persistent challenge in neural machine translation (NMT), particularly when they appear in discontinuous forms. In this context, the present study evaluates the ability of general-purpose large language models (LLMs) to address these limitations by systematically comparing the performance of Google Translate, DeepL, and GPT-4o in translating Spanish-to-English MWEs. A dataset of 600 examples—balanced between continuous and discontinuous MWEs—was machine translated and manually evaluated by two expert linguists. The results indicate that GPT-4o statistically significantly outperforms NMT systems in both forms of the MWEs, while DeepL and Google Translate exhibit substantial declines in performance for discontinuous MWEs. The results from this study show that general-purpose LLMs handle syntactic flexibility and idiomaticity more effectively than traditional NMT approaches. The study also contributes an open-source dataset and invites further research on LLM applications in MT.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

ChatGPT vs. DeepL vs. Google Translate: a human evaluation of multiword expressions’ machine translation quality — 科研速览 Science Skim