科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Academic radiology2026-08-31

Performance of Frontier Large Language Models on an FRCR Part 1-Style Examination: A Comparative Study of Physics and Anatomy Modules.

Hamza Eren Güzel, Başak Ünverdi

一句话结论 · In one sentence

In this custom FRCR Part 1-style assessment, large language models performed strongly on text-based physics questions but showed more variable performance on image-based anatomy, supporting their potential educational role while highlighting persistent limitations in precise anatomical localization.

原始摘要(英文原文)· Original abstract
RATIONALE AND OBJECTIVES: This study aims to compare the performance of three frontier large language models on an Fellowship of the Royal College of Radiologists (FRCR) Part 1-style examination comprising physics and image-based anatomy modules. MATERIALS AND METHODS: GPT-5.4 Thinking, Gemini 3.1 Pro, and Claude Opus 4.6 were evaluated using a custom FRCR Part 1-style examination. The physics module included 40 true/false questions comprising 200 independently scored statements, while the anatomy module included 100 image-based questions scored on a 0-2 scale. Both modules were structured according to the published FRCR Part 1 topic and modality distributions and were reviewed by two radiologists who had previously passed the examination. Model performance was compared using Cochran's Q and exact McNemar tests for physics and Friedman and Wilcoxon signed-rank tests for anatomy. RESULTS: Physics scores were 195/200 for Claude, 189/200 for GPT, and 187/200 for Gemini (overall P = 0.018). After Holm correction, Claude outperformed Gemini (adjusted P = 0.023), whereas the other pairwise differences were not significant. Anatomy scores were 160/200 for Gemini, 137/200 for GPT, and 83/200 for Claude (overall P < 0.001); all pairwise comparisons remained significant after Holm correction (adjusted P ≤ 0.038). CONCLUSION: In this custom FRCR Part 1-style assessment, large language models performed strongly on text-based physics questions but showed more variable performance on image-based anatomy, supporting their potential educational role while highlighting persistent limitations in precise anatomical localization.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Performance of Frontier Large Language Models on an FRCR Part 1-Style Examination: A Comparative Study of Physics and Anatomy Modules. — 科研速览 Science Skim