科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JDDG Journal der Deutschen Dermatologischen Gesellschaft2026-04-02· Interpretability

Beyond black‐box AI: Comparing ChatGPT‐4 interpretability and accuracy to CNNs in melanocytic lesions diagnosis

Ofer Reiter, Cristian Navarrete‐Dechent, Mor Atlas, Nir Nathansohn, Yaron Ben Mordehai, Tomer Mimouni, Romi Gleicher, Mahdi Awwad, I. Glenn Cohen, Ziad Khamaysi, Jonathan Shapiro

原始摘要(英文原文)· Original abstract
BACKGROUND: Artificial intelligence (AI) algorithms have advanced and recently shown high accuracy in diagnosing skin cancer from dermoscopic images. This study compared the diagnostic performance of the large language model ChatGPT-4 with that of specialized convolutional neural network (CNN)-based models in analyzing melanocytic lesions. PATIENTS AND METHODS: A cross-sectional comparative study was conducted using 117 dermoscopic images. The performance of ChatGPT-4 was assessed under two conditions: diagnosing lesions directly without annotations and diagnosing after annotating dermoscopic features. Results were compared with CNN-based models (YPSONO and ResNet) and human expert evaluations. The confusion matrices of all the models were calculated in addition to the diagnostic accuracy, sensitivity, specificity, and interobserver agreement (Cohen's Kappa). RESULTS: ChatGPT-4 achieved 92 % sensitivity, 89 % specificity, and an accuracy of 89.7 % in direct diagnosis. When annotations were required, sensitivity and specificity dropped to 68 % and 64 %, respectively. Agreement with experts on dermoscopic patterns was minimal (Cohen's Kappa = 0.13). ChatGPT-4 outperformed CNN models in direct diagnosis but exhibited notable limitations in describing dermoscopic features. CONCLUSIONS: ChatGPT-4 demonstrated promising potential for accurate melanoma versus nevus classification without annotations, surpassing CNN-based models. However, its limited ability to describe dermoscopic features accurately highlights the need for further research and training.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Beyond black‐box AI: Comparing ChatGPT‐4 interpretability and accuracy to CNNs in melanocytic lesions diagnosis — 科研速览 Science Skim