科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Big Data2025-12-15· Lagging

Benchmarking LLMs Against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

Jianhao Yan, Pingchuan Yan, Yulong Chen, Jing Li, Xianchao Zhu, Yue Zhang

原始摘要(英文原文)· Original abstract
This study presents a comprehensive evaluation of the translation capabilities of existing LLMs, such as GPT-4, ALMA-R, and Deepseek-R1, compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three language pairs (Chinese$\longleftrightarrow$English, Russian$\longleftrightarrow$English, and Chinese$\longleftrightarrow$Hindi) and three domains (News, Technology, and Biomedical). Our findings reveal that LLMs achieve performance comparable to junior-level translators in terms of total errors, while still lagging behind senior translators. Unlike traditional Neural Machine Translation systems, which show significant performance degradation in resource-poor language directions, LLMs like GPT-4 maintain consistent translation quality across all evaluated language pairs. Through qualitative analysis, we identify distinctive patterns in translation approaches: GPT-4 tends toward overly literal translations and exhibits lexical inconsistency, while human translators sometimes over-interpret context and introduce hallucinations. This study presents a systematic comparison between LLMs and human translators across different proficiency levels, providing valuable insights into the current capabilities and limitations of LLM-based translation systems.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking LLMs Against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels — 科研速览 Science Skim