科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Applied clinical informatics2026-08-28· Consistency (knowledge bases)

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency.

Moshe Carroll, Sabrina Kentis, Hannah Kareff, Clyde Schechter, Sunit Jariwala

一句话结论 · In one sentence

On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language models are rapidly transforming medical education, yet their performance in Allergy/Immunology remains insufficiently characterized. Furthermore, concerns regarding accuracy, consistency, and sensitivity to input format persist. OBJECTIVES: To evaluate and compare the accuracy and response consistency of three leading large language models-ChatGPT-5, Gemini 2.5, and Grok 4-on Allergy/Immunology United States Medical Licensing Examination Step 1-style questions under different prompt conditions. METHODS: Thirty-five United States Medical Licensing Examination Step 1-style questions were selected. Questions were presented to each model in two formats: single-question prompts and a combined prompt containing all questions. Fifteen trials were conducted for each format per model. Performance was assessed using mean accuracy and variability was measured using Shannon entropy. Mixed effects models tested effects of model, prompt condition, and question difficulty. RESULTS: Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) achieving higher mean scores than ChatGPT (74.3%). Single-item prompts yielded superior performance with Grok (93.1%) and Gemini (90.9%) demonstrating the highest accuracy. Transitioning to a combined prompt significantly reduced accuracy for all models. Accuracy also decreased with increasing question difficulty for all models. Grok demonstrated superior reliability, maintaining the lowest overall response entropy, whereas ChatGPT exhibited the highest variability. CONCLUSIONS: On Allergy/Immunology Step 1-style questions, Gemini and Grok demonstrated higher accuracy than ChatGPT, although their overall accuracies remained approximately 81%. Grok offered the most consistent performance. All models demonstrated substantial sensitivity to prompt complexity and inherent performance limitations. These findings underscore the importance of prompt optimization and support the supplementary role of these models in medical education.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency. — 科研速览 Science Skim