Hamza Eren Güzel, Başak Ünverdi
In this custom FRCR Part 1-style assessment, large language models performed strongly on text-based physics questions but showed more variable performance on image-based anatomy, supporting their potential educational role while highlighting persistent limitations in precise anatomical localization.
RATIONALE AND OBJECTIVES: This study aims to compare the performance of three frontier large language models on an Fellowship of the Royal College of Radiologists (FRCR) Part 1-style examination comprising physics and image-based anatomy modules.
MATERIALS AND METHODS: GPT-5.4 Thinking, Gemini 3.1 Pro, and Claude Opus 4.6 were evaluated using a custom FRCR Part 1-style examination. The physics module included 40 true/false questions comprising 200 independently scored statements, while the anatomy module included 100 image-based questions scored on a 0-2 scale. Both modules were structured according to the published FRCR Part 1 topic and modality distributions and were reviewed by two radiologists who had previously passed the examination. Model performance was compared using Cochran's Q and exact McNemar tests for physics and Friedman and Wilcoxon signed-rank tests for anatomy.
RESULTS: Physics scores were 195/200 for Claude, 189/200 for GPT, and 187/200 for Gemini (overall P = 0.018). After Holm correction, Claude outperformed Gemini (adjusted P = 0.023), whereas the other pairwise differences were not significant. Anatomy scores were 160/200 for Gemini, 137/200 for GPT, and 83/200 for Claude (overall P < 0.001); all pairwise comparisons remained significant after Holm correction (adjusted P ≤ 0.038).
CONCLUSION: In this custom FRCR Part 1-style assessment, large language models performed strongly on text-based physics questions but showed more variable performance on image-based anatomy, supporting their potential educational role while highlighting persistent limitations in precise anatomical localization.