Syeda Sitara Waseem, Saman Iftikhar, Ammar Rafiq, Kangyoon Lee, Syed Rizwan Hassan
Early detection of skin cancer significantly improves patient outcomes yet remains challenging due to the visual similarity of lesions and the need for clinical context. This paper proposes a novel Image-driven multimodal deep learning framework for automated skin cancer diagnosis integrating dermoscopic imaging sensors' clinical metadata and biomedical knowledge graphs. Dermoscopic images acquired from optical skin imaging sensors undergo advanced signal processing including multi-scale CLAHE enhancement and artifact suppression. A YOLOv8-nano sensor-based lesion detector achieves 97.2% mAP@0.5 for region-of-interest extraction. The biomedical knowledge graph (26 nodes 86 edges) encodes dermatological domain knowledge into 32-dimensional disease embeddings. Clinical metadata (age gender anatomical site) is encoded via learned feature vectors. A cross-attention fusion mechanism integrates these multimodal sensor-derived signals. Evaluated on the HAM10000 dataset (10,015 images 7 classes), our full multimodal model achieves 91.35% accuracy outperforming image-only (86.96%) and image+metadata (88.45%) baselines. An ensemble reaches 92.78% state-of-the-art accuracy with significant improvements on challenging classes (BCC +6.0% melanoma +3.8%). Statistical significance is confirmed (p < 0.01 95% CI [90.7% 92.0%]). This work demonstrates that Image-based multimodal signal fusion with structured medical knowledge enhances both diagnostic accuracy and clinical interpretability.