Fouad Abdulameer Salman, Bakhtawar Baluch, Zuriana Abu Bakar, Assane Lo
The rapid growth of scholarly publishing has increased the need of reliably identify credible and non-credible journal websites. Evaluating these entities requires close, cautious, thorough, and at times skeptical scrutiny by researchers, institutions, and funding agencies. However, existing automated detection approaches appear to be customized for specific aspects of the criteria and do not effectively detect non-credible entities. Therefore, in this paper, we come up with a comprehensive dataset comprising 4,174 manually verified journal websites. We classify the dataset into two categories: credible journals and non-credible journals, including predatory, commercial, and hijacked journals. We also represent 36 new features derived from different resources of the journal websites. We used these new features to train and test the proposed model to improve accuracy. XGBoost achieved a precision of 0.9778, recall of 0.9601, F1-score of 0.9689, and accuracy of 97.86%, outperforming both GBC and Random Forest in the experimental results. Feature-group experiments show that Web-Based System and Academic Integrity features substantially improve overall classification performance, while HTML-based characteristics provide strong discriminatory power. The proposed approach shows the potential of integrating technical, structural, and scholarly-verification signals to support automated journal-credibility assessment and give researchers an efficient way to identify potentially non-credible scholarly venues.