Luigi A Vaira, Jerome R Lechien, Fabiola Giudici, Andrea Gasparato, Antonino Maniaci, Andrea De Vito, Fabio Maglitto, Miguel Mayo-Yáñez, Stefania Troise, Carlos M Chiesa-Estomba, Guido Gabriele, Andrea Frosolini, Thomas Radulesco, Fabiana Allevi, Valentino Vellone, Umberto Committeri, Fabrizio Spallaccia, Resi Pucci, Giuseppe Consorti, Giulio Cirignaco, Salvatore Crimi, Marzia Petrocelli, Eleonora M C Trecca, Giannicola Iannella, Francesco Laganà, Bernardo Bianchi, Alberto Maria Saibene, Sholem Hack, Paolo Boscolo-Rizzo, Giovanni M Soro, Giovanni Salzano, Jerry Polesel, Giacomo De Riu
In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. However, given the exploratory single-reader design, these findings should not be interpreted as evidence of equivalence or superiority relative to clinicians and require further prospective external validation before any potential clinical deployment in telemedicine or primary care settings.
OBJECTIVE: To evaluate the real-world diagnostic performance of a multimodal large language model (LLM) for image-based assessment of oral mucosal lesions compared with clinicians of varying expertise.
STUDY DESIGN: Prospective international multicenter diagnostic accuracy study.
SETTING: Twenty university and tertiary head and neck centers in Italy, Belgium, France, Spain, and Israel.
METHODS: We enrolled 350 consecutive patients (320 with oral lesions, 30 with normal mucosa). Clinical photographs and basic epidemiologic data were analyzed using Gemini 2.5 Advanced with a standardized prompt. Model outputs for lesion detection, malignancy versus benign versus normal, precise histologic diagnosis, and urgency class were compared with histopathology and with 4 clinicians. Sensitivity, specificity, accuracy, and agreement were calculated.
RESULTS: AI-Gemini achieved 97.1% accuracy for lesion detection and malignancy classification, with sensitivity 98.5% and specificity 96.2% for malignancy, and 88.0% accuracy for precise histologic diagnosis. The head and neck surgeon achieved the highest accuracy for precise diagnosis (97.7%). Three-class diagnostic accuracy was 94.2% for AI-Gemini and 67.0% to 86.5% for nonsurgeon clinicians. Urgency assignment was correct in 70% of cases (κ = 0.716), with a conservative tendency to overestimate risk. Agreement with histologic diagnosis was almost perfect (κ = 0.929).
CONCLUSION: In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. However, given the exploratory single-reader design, these findings should not be interpreted as evidence of equivalence or superiority relative to clinicians and require further prospective external validation before any potential clinical deployment in telemedicine or primary care settings.