科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Academic radiology2026-08-12

Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large Language Models.

Matteo Haupt, Martin H Maurer

一句话结论 · In one sentence

Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.

原始摘要(英文原文)· Original abstract
RATIONALE AND OBJECTIVES: Multimodal large language models (LLMs) are increasingly applied to image-based radiology tasks, but their diagnostic accuracy across clinically distinct populations remains poorly characterized. We quantified age-related differences in zero-shot LLM performance for pneumonia detection on pediatric vs. adult chest radiographs and compared generalization gaps with a domain-trained convolutional neural network (CNN) baseline. MATERIALS AND METHODS: GPT-5.2 (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Pro (Google) were evaluated zero-shot on balanced pediatric and adult test sets of frontal chest radiographs (n = 1000 each; 500 pneumonia/500 normal). Cohort-specific InceptionV3 CNNs were trained on the remaining development-pool images (pediatric n = 4715; adult n = 13,863) and evaluated on the same test sets. Performance was assessed using the Matthews correlation coefficient (MCC) with 95% bootstrap confidence intervals (CIs); domain shift was quantified as Δ = Adult - Pediatric. RESULTS: In pediatrics, the CNN outperformed all LLMs (MCC 0.799, 95% CI 0.766-0.832) vs. GPT-5.2 (0.484, 0.436-0.532), Claude Opus 4.5 (0.470, 0.418-0.521), and Gemini 2.5 Pro (0.272, 0.224-0.316). In adults, all models improved, but the CNN remained best (MCC 0.850, 0.816-0.882). Age-related gains were larger for LLMs (ΔMCC +0.220 [95% CI 0.160-0.281] to +0.466 [0.407-0.525]) than for the CNN (ΔMCC +0.051 [0.005-0.098]), driven mainly by specificity increases. CONCLUSION: Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large Language Models. — 科研速览 Science Skim