科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Indian Association of Pediatric Surgeons2026-01-01

Diagnostic Performance of Multimodal Large Language Models for Grading Pediatric Vesicoureteral Reflux on MCUG: A Preliminary Study.

Shubham Saini, Anmol Bhatia, Akshay Kumar Saxena, Ravi Prakash Kanojia, Kushaljit Singh Sodhi

一句话结论 · In one sentence

ChatGPT-5 and Gemini 2.5 Pro demonstrated fair but suboptimal accuracy in grading VUR. Model performance appears to be influenced by the prompting strategy used. Multimodal large language models show potential as adjunctive tools; however, further optimization of prompting approaches and multimodal integration is required before clinical application.

原始摘要(英文原文)· Original abstract
AIMS: To evaluate the diagnostic accuracy of ChatGPT-5 and Gemini 2.5 Pro compared with a pediatric radiologist reference standard in grading vesicoureteral reflux (VUR) on pediatric micturating cystourethrogram (MCUG) studies. MATERIALS AND METHODS: This retrospective study included 125 pediatric MCUG cases, comprising 25 cases per VUR Grade (I-V) as determined by radiologists. A balanced distribution of VUR grades was employed to facilitate a comprehensive assessment of model performance across the severity spectrum. For each case, the most representative image was anonymized, and the stored Joint Photographic Experts Group images were randomly uploaded to ChatGPT-5 and Gemini 2.5 Pro with a standardized prompt requesting VUR grading. Grades assigned by the models were compared with the radiologist's consensus standard. RESULTS: VUR was present in 182/250 (72.8%) renal units, more frequently on the left side (80%) than the right side (65.6%). Both models demonstrated the highest accuracy in identifying the absence of reflux (ChatGPT-5 = 92.6% and Gemini 2.5 Pro = 88.2%). Overall diagnostic accuracy was 52% (131/250) for ChatGPT-5 and 46% (115/250) for Gemini 2.5 Pro. ChatGPT-5 performed best for grades III-IV (51% and 49%, respectively), whereas Gemini 2.5 Pro performed best for Grade V (54%). Agreement with radiologists was fair (ChatGPT-5 κ = 0.40; Gemini 2.5 Pro κ = 0.33). CONCLUSION: ChatGPT-5 and Gemini 2.5 Pro demonstrated fair but suboptimal accuracy in grading VUR. Model performance appears to be influenced by the prompting strategy used. Multimodal large language models show potential as adjunctive tools; however, further optimization of prompting approaches and multimodal integration is required before clinical application.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diagnostic Performance of Multimodal Large Language Models for Grading Pediatric Vesicoureteral Reflux on MCUG: A Preliminary Study. — 科研速览 Science Skim