Hadeel Saed
Social-media sentiment classification is difficult because posts are short, informal, context-dependent, and unevenly distributed across classes. It was observed that five TF-IDF classifiers (Logistic Regression, Linear SVM, Multinomial Naïve Bayes, Random Forest, and Gradient Boosting) were compared with the compact DistilBERT and MiniLM models. Fixed TweetEval consisted of 45,615 training posts, 2,000 validation posts, and 12,284 test posts. The main selection criterion was the validation macro-F1. To sum up, MiniLM outperformed the strongest classical baseline, Logistic Regression, by 11.8573 macro-F1 points on test set. This improvement was confirmed by Holm-adjusted McNemar and bootstrap analyses and was considered to be statistically and practically significant. Even though it required more resources than classical alternatives, MiniLM met the required criteria predefined, and in addition had a smaller artifact and lower energy proxy than DistilBERT. Explainability used global Logistic Regression TF-IDF coefficients and local DistilBERT Integrated Gradients. Overall, top-token masking reduced the probability more than the random masking in 30 cases, lending support to the attribution faithfulness. Findings help in the selection of practical opinion-monitoring models.