科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Diagnostics (Basel, Switzerland)2026-08-28

Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering Diseases: A Multicenter Paired Study.

Asmaa Abou-Bakr, Salma M Saad, Nevine H Kheir El Din, Abdullah Bin Nabhan, Amal Bajonaid, Asma Saleh Almeslet, Fatma E A Hassanein

原始摘要(英文原文)· Original abstract
Background/Objectives: To compare the diagnostic performance of Claude Opus 4.7 and Gemini Pro 3 for the differential diagnosis of oral autoimmune blistering diseases (AIBDs) and evaluate their diagnostic reasoning, confidence, and calibration. Materials and Methods: This retrospective multicenter paired diagnostic accuracy study included 200 clinicopathologically confirmed AIBD cases (50 each of pemphigus vulgaris, mucous membrane pemphigoid, bullous pemphigoid, and linear IgA bullous dermatosis). Each case was independently assessed by both models using identical standardized clinical information and clinical photographic inputs. The task required forced-choice classification among the four predefined diseases. Histopathological and direct immunofluorescence findings were used exclusively to establish the clinicopathological reference diagnosis and were not provided to the AI models. The reference diagnosis was established by clinicopathological correlation. The primary outcome was diagnostic accuracy. Secondary outcomes included disease-specific diagnostic performance, Cohen's κ, ROC analysis, calibration, confidence, reasoning quality, management recommendations, and error patterns. Pre-consensus inter-rater reliability of the two human assessors was also evaluated using Cohen's κ for binary outcomes and weighted Cohen's κ for the ordinal reasoning-quality score. Results: Claude achieved significantly higher diagnostic accuracy than Gemini (92.0% vs. 86.0%, p = 0.012), stronger agreement with the reference standard (κ = 0.893 vs. 0.813), and superior discrimination (macro-AUC 0.998 vs. 0.965). Claude demonstrated higher key diagnostic-feature identification (92.0% vs. 86.0%; p = 0.012) and higher clinical-reasoning scores (61.0% vs. 40.0% of responses rated good; Wilcoxon p < 0.001; r = 0.47), whereas management recommendations did not differ significantly (100.0% vs. 98.0%; p = 0.125). Calibration results were metric-dependent: Claude had a lower one-vs-rest Brier score (0.0468 vs. 0.0558), whereas Gemini had a lower expected calibration error (0.083 vs. 0.251). For both models, the predominant error was misclassification of linear IgA bullous dermatosis as mucous membrane pemphigoid. Conclusions: Both multimodal LLMs showed high performance in this controlled four-class benchmark, with Claude Opus 4.7 outperforming Gemini Pro 3 in overall accuracy and reasoning quality. However, these findings do not establish autonomous diagnostic capability, clinical effectiveness, or safety. The LABD-MMP misclassification and metric-dependent calibration highlight important limitations. The models should therefore be regarded as investigational adjunctive decision-support tools requiring clinician oversight and diagnostic verification. Prospective external and human-in-the-loop validation is required before clinical implementation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering Diseases: A Multicenter Paired Study. — 科研速览 Science Skim