Bhagyashri S. Sonune, R Udaykumar, Shon G. Nemane, Dhiraj P. Tulaskar, Manish Bhaiyya, Hossam Haick
The problem of accurate and interpretable automated classification of dermoscopic skin lesions is difficult because publicly available benchmark datasets are imbalanced, diagnostically heterogeneous, and prone to image-level confounding effects. While individual techniques such as pre-trained convolutional neural networks (CNNs), ensemble learning, and explainable artificial intelligence (XAI) are established, their combined effectiveness is reported without controlled comparison to determine whether ensemble consensus improves robustness and explanation quality in a unified framework. In this paper, we report a controlled comparison of six pre-trained CNN architectures, namely ResNet50, Xception, EfficientNetB3, MobileNetV2, DenseNet201, and InceptionV3, on the publicly available HAM10000 dataset, as well as their soft voting ensemble. The contribution of this paper is not in proposing a new deep learning architecture but rather in presenting a harmonized evaluation framework to determine whether diversity in architectures can be exploited to improve robustness on minority classes and qualitatively better explanations in terms of spatial coherence in seven-class classification of dermoscopic images of skin lesions. To this end, we used Grad-CAM, LIME, and occlusion sensitivity methods to evaluate the interpretability of each individual CNN as well as the ensemble. The soft voting ensemble outperformed all individual architectures by achieving a macro-average ROC curve of 0.985 and accuracy of 89.37%, while simultaneously providing qualitatively better spatially coherent explanations that are lesion-centric, suggesting better trustworthiness in transparent XAI-based automated screening in dermatological studies.