Pablo Antonio López-Galindo, Giuseppe Troiano, Matteo Serroni, Debora R Dias, Pasquale Santamaría, Luigi Nibali, Gustavo Avila-Ortiz, Andrea Ravidà
Clinicians outperformed AI under a fixed threshold, but AI performance was comparable to clinicians using optimal thresholds, situating AI within the range of experienced clinicians despite modest overall predictive ability.
AIM: To compare the prognostic performance of an artificial intelligence (AI) model with that of experienced clinicians in predicting tooth loss over a 10-year period.
MATERIALS AND METHODS: An AI model trained on structured clinical and radiographic data was compared with 12 periodontists and 11 general dentists (GDs), who independently assigned prognostic scores (0-10 scale) to 300 teeth with known 10-year outcomes. AI and clinician performance were evaluated under a fixed threshold and group-specific optimal thresholds. Accuracy, sensitivity, specificity, predictive values, area under the receiver operating characteristic curve and calibration were used for comparison.
RESULTS: Using the fixed threshold (score > 5 = survival), clinicians achieved higher overall accuracy than AI (75.6% periodontists, 74.9% GDs, 69.2% AI; p < 0.05), with sensitivity low and comparable across groups (14.7%-22.7%). Using group-specific optimal thresholds, AI achieved the highest accuracy for all-cause tooth loss (67.5%) and performed comparably to GDs for periodontitis-related tooth loss (76.2% vs. 76.3%). Calibration was substantially worse for AI (Brier score 0.247) than for periodontists and GDs (0.171, 0.170). Individually, AI performance overlapped with that of several clinicians.
CONCLUSION: Clinicians outperformed AI under a fixed threshold, but AI performance was comparable to clinicians using optimal thresholds, situating AI within the range of experienced clinicians despite modest overall predictive ability.