科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Children (Basel, Switzerland)2026-09-18

Comparative Expert Evaluation of Multimodal Large Language Models for Pediatric Rash Diagnosis: Clinical Utility, Safety, Information Quality, and Readability.

Dilara Lahut, Özlem Erdede, Rabia Gönül Sezer Yamanel

原始摘要(英文原文)· Original abstract
Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods: Fifteen content-validated pediatric rash vignettes with brief histories and anonymized photographs were submitted once to each platform using a standardized zero-shot prompt. Three pediatricians blinded to platform identity independently rated the 45 responses using a five-point Clinical Utility and Safety (CUS) scale and a five-item modified DISCERN instrument. Diagnostic correctness was assessed descriptively; platform comparisons used Friedman tests with Bonferroni-adjusted Wilcoxon tests when appropriate. Results: Overall, 82.2% of CUS ratings were in categories 4-5 and 83.0% of modified DISCERN scores were ≥20/25; no rating was assigned to CUS category 1. Gemini and Grok had descriptively higher expert ratings than ChatGPT, but CUS did not differ significantly across platforms (p = 0.157), and although modified DISCERN differed globally (p = 0.038), no pairwise comparison remained significant after adjustment. In the single-query diagnostic assessment, at least one platform missed the reference diagnosis in 9/15 vignettes, and all three missed porphyria. Gemini generated the longest responses, whereas Grok produced the most linguistically complex text; neither response length nor readability was associated with expert ratings. Conclusions: The three multimodal LLMs produced predominantly clinically acceptable responses, but performance varied by vignette and platform. Because each vignette-platform combination was sampled once, diagnostic findings represent single-response observations rather than stable platform accuracy estimates. Clinical verification remains necessary.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Comparative Expert Evaluation of Multimodal Large Language Models for Pediatric Rash Diagnosis: Clinical Utility, Safety, Information Quality, and Readability. — 科研速览 Science Skim