Goran Mitreski, Philip Chan, Cephas Tamburayi Kamba, Nathan Ng, David Duong, Rose Thomas, Jyothirmayi Velaga, Miranda Siemienowicz, Hong Kuan Kok
GPT-4o achieved high translation quality across multiple interventional radiology consent simulations, demonstrating the feasibility of real-time artificial intelligence interpretation for many languages. However, performance was language-dependent; certain languages (Cantonese and Shona) remained sub-optimal and merit human-interpreter supervision. These findings represent a platform-specific proof-of-concept evaluation rather than validation of artificial intelligence-based medical translation more broadly.
PURPOSE: Language barriers in healthcare lead to miscommunication, reduced comprehension, and adverse outcomes. Professional medical interpreters are the gold standard, but are often unavailable for interventional radiology consent discussions. Recent advances in large language models such as GPT-4o have shown high accuracy in multilingual tasks. We conducted a proof-of-concept internal validation of GPT-4o as a live bilingual translation tool in interventional radiology consent simulations across nine languages.
MATERIAL AND METHODS: Thirty-six simulated consent sessions were conducted across nine languages - Shona, Macedonian, Polish, Tamil, Cantonese, Mandarin, Malay (Bahasa Malaysia), Afrikaans, and Vietnamese, with four procedural scenarios evaluated per language (Transjugular Intrahepatic Portosystemic Shunt, Yttrium-90 Radioembolization, liver biopsy, large-bore mechanical thrombectomy) using GPT-4o to perform real-time translation between English and native language speaking Interventional Radiologists/Diagnostic Radiologists. A 13-domain rubric of translation adequacy (each scored 1-5) produced a maximum possible total of 65 points. Sessions scoring ≥50/65 were considered acceptable.
RESULTS: Across 36 sessions, mean scores were 56.2 ± 6.3 (86.5% of maximum). Acceptability (≥50) was achieved in 31/36 sessions (86.1%). Seven languages (Mandarin, Macedonian, Polish, Tamil, Malay, Afrikaans, and Vietnamese) consistently exceeded the threshold, while Cantonese (mean ~40/65) and Shona (mean ~42/65) lagged. Translation speed (time to output) averaged 4.8/5 across all languages.
CONCLUSIONS: GPT-4o achieved high translation quality across multiple interventional radiology consent simulations, demonstrating the feasibility of real-time artificial intelligence interpretation for many languages. However, performance was language-dependent; certain languages (Cantonese and Shona) remained sub-optimal and merit human-interpreter supervision. These findings represent a platform-specific proof-of-concept evaluation rather than validation of artificial intelligence-based medical translation more broadly.