Oluwaseyi Oladejo, Ahmed Abdelmoamen Ahmed
The exponential growth of cyber threats in modern digital infrastructure demands advanced detection systems that adapt to evolving attack patterns. Traditional cybersecurity approaches struggle with dynamic threats, requiring extensive labeled datasets and retraining for each new category. This paper presents a comprehensive transfer learning framework for cybersecurity threat detection, leveraging the CICIoMT dataset as a benchmark to enhance detection capabilities across heterogeneous cybersecurity environments. We propose a machine learning (ML)-enabled framework that employs systematic feature alignment, hybrid class balancing, and multi-algorithm evaluation using machine learning models, including Multi-Layer Perceptron (MLP), Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting, and XGBoost. The proposed approach addresses the critical challenges of data scarcity and domain heterogeneity in cybersecurity by enhancing feature engineering with cybersecurity-specific features, statistical aggregations, and PCA embeddings. Extensive experimental evaluation across two target datasets (CICIoT and IoT-23) demonstrates both the exceptional successes and critical limitations of cross-domain transfer learning in cybersecurity. The framework achieved outstanding performance on domain-compatible datasets, with RF reaching 99.0% accuracy on CICIoT, Gradient Boosting achieving 98.9%, and XGBoost delivering 98.4%, demonstrating exceptional knowledge transfer from medical IoT to smart home IoT environments. However, transfer learning to IoT-23 was unsuccessful (50% accuracy, equivalent to random guessing), revealing that feature domain difference, where identical attack labels encode fundamentally different behavioral patterns, prevents effective knowledge transfer despite nominal class overlap. This research makes significant advances in adaptive cybersecurity systems by providing a rigorous evaluation of both the successes and limitations of transfer learning. This work demonstrates that ensemble methods (RF, XGBoost, and Gradient Boosting) achieve superior cross-domain performance compared with neural networks on compatible domains, while also revealing fundamental challenges when the source and target domains differ in their feature spaces.