Hiruna K Diyasena, Aadam Ahmad, Thenuk H Weerasinghe, Olivia V Crawshaw, Anna A O'Brien
Introduction Large language models (LLMs), such as Generative Pre-trained Transformer (GPT)-5.2 (OpenAI, Inc., San Francisco, United States), are increasingly used as revision tools for postgraduate medical examinations, including the Membership of the Royal College of Physicians (MRCP) Part 1 assessment. The educational value of newer extended-reasoning modes within the same model remains largely unexplored in the literature, and it is unclear whether improvements in performance or changes in the nature and frequency of errors can justify their greater computational cost. Materials and methods A paired design was used on a sample paper comprising 193 MRCP Part 1 single-best-answer (SBA) questions, with each question submitted three times to each mode: low-latency GPT-5.2 Instant and deeper-reasoning GPT-5.2 Thinking. The primary analysis compared the two modes at the question level. For each mode, a question was classified as majority correct when at least two of its three trials were correct, and the paired majority outcomes were analyzed using McNemar's test. A secondary trial-level analysis used a logistic generalized estimating equation (GEE) model to account for repeated observations within questions. Incorrect responses were isolated and independently categorized by two reviewers blinded to mode using an adapted taxonomy of cognitive bias-like errors. Owing to sparse error counts, the findings were analyzed descriptively. Results Instant answered 549/579 correctly (94.8%), and Thinking answered 563/579 correctly (97.2%). Out of 193 questions, both modes achieved majority-correct responses in 183 questions (94.8%) and majority-incorrect responses in three questions (1.6%). There was a total of seven discordant pairs. Discordant-pair analysis found no statistically significant difference between modes (odds ratio (OR) = 6.0; 95% CI (0.72 - 49.84); P = 0.125), whereas trial-level analysis adjusted for operating mode alone showed higher odds of a correct response in favor of Thinking mode (OR = 1.92, 95% CI (1.14 - 3.24), P = 0.014). Instant mode errors were concentrated in input-conflicting hallucination and availability bias, whereas Thinking mode errors were more evenly distributed and included a greater relative proportion of reasoning hallucination. Conclusion Both GPT-5.2 reasoning modes achieved high and comparable performance in MRCP Part 1 SBA questions. Thinking mode was marginally more consistent across trials, but ceiling effects limited any practical benefit. Hence, low-latency modes may offer comparable practical utility in this structured, recall-dominant assessment context. Thinking mode demonstrated different error patterns when compared to Instant mode, but did not eliminate errors entirely; learners should critically appraise all model outputs.