科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Journal of Infectious Diseases2026-05-07· Triage

Diagnostic Accuracy of Commercial Large Language Models for Anogenital Skin Lesion Images: A Comparative Study of Gemini, Claude, and ChatGPT

Nyi Nyi Soe, Phyu Mon Latt, David Lee, Ei T. Aung, Ryan Horn, Jason J. Ong, Christopher K Fairley, Eric P F Chow

原始摘要(英文原文)· Original abstract
BACKGROUND: Diagnosing anogenital dermatological conditions often requires specialist expertise that is unavailable in many clinical settings. Large language models (LLMs) are increasingly accessible to clinicians, but their diagnostic accuracy for anogenital dermatology has not been evaluated. We evaluated the diagnostic accuracy of three LLMs (Gemini 2.5 Pro, Claude Opus 4.1, and ChatGPT 5 Thinking). METHODS: This study was conducted between September and November 2025, using de-identified clinical images of anogenital conditions from the STI Atlas (stiatlas.org (https://stiatlas.org/)) and other publicly available sources. Primary outcomes were correct classification of images identified as sexually transmitted infections (STIs) vs non-STIs and the inclusion of the correct diagnosis among the LLMs' top-ranked (top-1), top-3, or top-5 differential diagnoses. RESULTS: Among 218 images, Gemini achieved the highest accuracy for STI binary classification (76.2% [95% CI, 70.5% - 81.9%]) and differential diagnosis (top-1, 39.0% [95% CI, 32.7% - 45.7%]; top-3, 54.6% [95% CI, 47.9% - 61.1%]; top-5, 60.6% [95% CI, 53.9% - 66.9%]), followed by ChatGPT and Claude. In subgroup analysis, all LLMs showed substantially reduced accuracy for diagnostically challenging images (top-5 accuracy range, 29.2% - 40.0%). Gemini consistently outperformed Claude across most subgroups (P < 0.05). None of the LLMs could identify any mpox correctly. CONCLUSION: LLMs showed limited accuracy for diagnosing anogenital conditions, particularly for challenging images. The best-performing model achieved only 39.0% for top-1 diagnosis, indicating that current LLMs cannot reliably diagnose anogenital conditions. These tools may support supervised clinical triage but need further validation before routine clinical use.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diagnostic Accuracy of Commercial Large Language Models for Anogenital Skin Lesion Images: A Comparative Study of Gemini, Claude, and ChatGPT — 科研速览 Science Skim