科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Access2026-01-01· Computer science

XGBoost-Based URL Phishing Detection Method With Cross-Dataset Validation

Milosz Misiek, Tomasz Hyla

原始摘要(英文原文)· Original abstract
This article analyses the performance characteristics of XGBoost models across multiple datasets for phishing URL detection, extending our previous conference paper with comprehensive cross-dataset validation and comparative analysis. Using a validation framework with three distinct datasets — a custom phishing dataset (75,738 samples), the large-scale GramBeddings dataset (639,723 samples), and the PhiUSIIL dataset (47,103 samples)—we demonstrate that XGBoost delivers robust performance across diverse data sources. Our approach addresses the issue of feature instability, showing how balanced feature engineering is crucial for reliable detection. The model achieves 90.7% accuracy and 91.2% F1 score on the custom dataset, with a solid 77.8% average accuracy across three independent test sets. In comparative benchmarks against Neural Networks, XGBoost proved superior in training efficiency and detection quality, achieving 2.1% higher accuracy and training 12.6 times faster (0.95s vs 12.01s) with 3.6 times less memory. Although Neural Networks offered faster inference (7.51ms vs 20.6ms) and smaller model sizes, XGBoost’s balance of high accuracy and rapid training makes it highly effective for practical, real-world phishing detection systems.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

XGBoost-Based URL Phishing Detection Method With Cross-Dataset Validation — 科研速览 Science Skim