科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Digital health2026-01-01

Assessing the accuracy and readability of large language models in answering Turkish-language patient-generated queries on radiological contrast agents.

Abdullah Enes Ataş

一句话结论 · In one sentence

LLMs generate highly readable and subjectively reassuring Turkish-language texts in response to patient queries regarding radiological contrast agents. However, due to domain-specific structural gaps regarding quantitative thresholds and lack of citations, they should be positioned as hybrid communication tools subject to mandatory review by specialist physicians rather than standalone medical advisors.

原始摘要(英文原文)· Original abstract
OBJECTIVE: To assess the clinical accuracy, structural readability, and Ensuring Quality Information for Patients (EQIP-36) quality of Large Language Model (LLM) responses to real patient queries regarding radiological contrast agents and to define their safe boundaries in practice. METHODS: The 30 most popular Turkish patient questions reflecting fears and risk perceptions about contrast agents were identified via Google Trends and AlsoAsked.com. These queries were presented to ChatGPT-5.4, Gemini 3, and Claude 4.6 Sonnet using a zero-shot learning approach. Responses were evaluated by experts for clinical accuracy against the 2024/2025 American College of Radiology (ACR) Manual on Contrast Media. Information quality and readability were measured using the expanded EQIP-36 scale and the Ateşman and Bezirci-Yılmaz indices. RESULTS: Claude 4.6 Sonnet demonstrated the highest clinical accuracy and high compliance with guidelines (>90%). Gemini 3 exhibited a more conservative stance, occasionally resulting in over-triage by exaggerating risks. While all models scored exceptionally well in the Content and Structure dimensions of the EQIP-36, generating readable texts, they universally failed in the Identification dimension by omitting author names, update dates, or bibliographic sources. Readability indices indicated that comprehending the texts required an average of 10.5 to 11.6 years of formal education. CONCLUSION: LLMs generate highly readable and subjectively reassuring Turkish-language texts in response to patient queries regarding radiological contrast agents. However, due to domain-specific structural gaps regarding quantitative thresholds and lack of citations, they should be positioned as hybrid communication tools subject to mandatory review by specialist physicians rather than standalone medical advisors.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Assessing the accuracy and readability of large language models in answering Turkish-language patient-generated queries on radiological contrast agents. — 科研速览 Science Skim