Charles Goodmaker, Rishi Bhandari, Anwar Tappuni, Tuan D Pham
Classifying oral cancer and OPMD lesions with visible light photographs is an expanding field of research, but one which reveals space for more robust and clinically grounded training datasets as well as exploration of frontier AI methodologies.
OBJECTIVES: To map the scope, methodological characteristics, and clinical task design of artificial intelligence (AI) models applied to visible-light photographs of the oral cavity for classification of oral cancer and oral potentially malignant disorders (OPMD).
METHODS: A scoping review following the Arksey and O'Malley framework and PRISMA-ScR was conducted. PubMed, Embase, Scopus, and Web of Science were searched from January 2015 to October 2025. Studies were eligible if they applied AI to intraoral photography for lesion detection, classification, or risk stratification. Data were extracted on dataset provenance, ground-truth labelling, model architecture, validation strategy, performance metrics, and reporting completeness.
RESULTS: 127 studies met inclusion and were categorised by their output classes across a mutually exclusive, five-category framework (four ordered tiers plus a residual category) with a benign-OPMD-malignant disease spectrum. Over half addressed a cancer-versus-non-cancer endpoint only (54.3%, n=69), whilst 27.6% (n=35) preserved OPMD as a discrete output category. Public image datasets were used in 44.1% (n=56) of studies, and fully histologically-confirmed ground-truth labels were available in only 7.9% (n=10). External validation was reported in 8.7% (n=11), patient-level data splitting in 10.2% (n=13), and prospective clinical evaluation in 0% (n=0). ResNet-family architectures were most frequently reported (50.4%, n=64, as primary or comparator model); explainable-AI components appeared in 37.0% (n=47) and multimodal inputs in 12.6% (n=16).
CONCLUSION: Classifying oral cancer and OPMD lesions with visible light photographs is an expanding field of research, but one which reveals space for more robust and clinically grounded training datasets as well as exploration of frontier AI methodologies.