Lorena Adriana Paun, Iulian Alexandru Taciuc, Mihai Dumitru, Daniela Vrinceanu, Andreea Marinescu, Alexandru-Darius Dragomir-Serboiu, Alina Oancea, Monica-Mihaela Cirstoiu, Adrian Costache
Background: Convolutional neural networks (CNNs) perform well in histopathological classification of oral squamous cell carcinoma (OSCC), but their robustness across distinct tissue domains remains insufficiently studied. This study assessed whether a CNN trained only on primary oral and oropharyngeal squamous cell carcinoma images remained transferable to an independent metastatic lymph node histopathology dataset. Methods: Three public datasets containing 14,760 OSCC/OPSCC and 8530 normal oral mucosa images were combined for model development. An ImageNet-pretrained EfficientNetB0 backbone was used as a fixed feature extractor with a task-specific binary classification head. Performance was first assessed on an independent internal testing subset and subsequently evaluated on 20,000 H&E-stained normal and metastatic lymph node patches from a separate public dataset. Results: The model achieved an internal testing accuracy of 91.48%, with 92.95% sensitivity, 88.92% specificity, 93.56% precision, a 93.25% F1-score, and a Youden's J index of 0.819. External validation resulted in a substantial decrease in overall classification performance, with an accuracy of 56.91%, specificity of 33.32%, precision of 48.71%, F1-score of 63.37%, balanced accuracy of 61.99%, MCC of 0.279, and a Youden's J index of 0.240. Nevertheless, sensitivity for metastatic tissue remained high at 90.66%, indicating a markedly asymmetric external error profile characterized predominantly by false-positive classifications. Conclusions: The marked performance decrease during external validation demonstrates the limitations of direct cross-domain transfer between substantially different histopathological environments. However, the preserved sensitivity suggests that some discriminative information remained transferable beyond the development domain. These findings support partial rather than universal cross-domain generalization and emphasize the importance of independent out-of-distribution evaluation when assessing deep learning robustness in computational pathology.