L Birkett, T Fowler, A Mazzoleni
Large language models are increasingly being used in medical education, but their performance in high-stakes anaesthesia examinations is not well defined. This study evaluated ChatGPT-5 on the official sample questions for the Fellowship of the Royal College of Anaesthetists Final Written examination. ChatGPT-5 was tested using questions from the 'Guide to the Fellowship of the Royal College of Anaesthetists Examination: The Final' publication. They comprised 100 marks of Constructed Response Questions, 900 Multiple True False stems and 60 Single Best Answer questions. Performance was compared with the published pass mark for Constructed Response Questions (62/100) and historical pass ranges for multiple true false (72-75%) and single best answer questions (55-60%). Subgroup analyses were performed by topic. A memorisation effects Levenshtein detector analysis assessed whether answers were likely present in the model's training data. ChatGPT-5 scored 89/100 in constructed response questions, 746/900 in multiple true false questions, and 39/60 in single best answer questions. All scores were above the pass standards. Physiology constructed response questions scored 12/12, higher than the overall constructed response question mean. Regional anaesthesia multiple true false questions scored 42/60, lower than the overall multiple true false mean. ChatGPT-5 achieved passing-level performance across all formats in the sample set. Performance was strongest in constructed response questions and weaker in single best answer questions. Large language models may therefore aid postgraduate revision and exam preparation.