Changez Khan, Awais Adnan, Sajid Anwar
The widespread use of online reviews has significantly influenced decision-making across various sectors. However, the presence of fake reviews estimated to comprise over 30% of all reviews, threatens consumer trust and platform reliability. This study focuses on detecting fake reviews by utilizing sentiment analysis and a range of linguistic features, including review length, punctuation, misspellings, adjective usage, and textual similarity. We utilize three publicly available datasets from diverse domains to ensure robustness and cross-domain applicability: 288 065 reviews from 395 open-source Android applications, 10 000 hotel reviews from the travel industry, and 24 386 product reviews from an e-commerce clothing platform. These datasets collectively offer a comprehensive view of user-generated content across mobile apps, hospitality, and online retail. The study employed advanced machine learning algorithms, including random forest (RF) and decision tree (DT) models, supplemented by techniques such as information gain and upper/lower approximation. These methods were applied to assess feature importance and classify reviews as original or fake. Results demonstrate that integrating linguistic features with sentiment analysis significantly improves detection performance. RF and DT classifiers achieved up to 98% accuracy and high F1 scores across all datasets. To interpret model predictions, SHapley Additive exPlanations analysis was employed, offering insights into the contribution of individual features to classification outcomes. While the current approach is effective, reliance on linguistic and sentiment-based indicators may be insufficient for detecting artificial intelligence (AI)-generated fake content. This study contributes to strengthening digital trust by offering a reliable, interpretable, and domain-independent method for fake reviews detection.