Carlos Manuel Hidalgo Ternero, Vicent Briva-Iglesias
Multiword expressions (MWEs) remain a persistent challenge in neural machine translation (NMT), particularly when they appear in discontinuous forms. In this context, the present study evaluates the ability of general-purpose large language models (LLMs) to address these limitations by systematically comparing the performance of Google Translate, DeepL, and GPT-4o in translating Spanish-to-English MWEs. A dataset of 600 examples—balanced between continuous and discontinuous MWEs—was machine translated and manually evaluated by two expert linguists. The results indicate that GPT-4o statistically significantly outperforms NMT systems in both forms of the MWEs, while DeepL and Google Translate exhibit substantial declines in performance for discontinuous MWEs. The results from this study show that general-purpose LLMs handle syntactic flexibility and idiomaticity more effectively than traditional NMT approaches. The study also contributes an open-source dataset and invites further research on LLM applications in MT.