Dina Shehada, Hissam Tawfik, Ahmed Bouridane, Abir Hussain
The continuous analysis of emotional cues through facial emotion recognition (FER) systems can support mental health evaluation and psychological well-being monitoring systems. Most FER systems face privacy and trust concerns due to their centralized data approaches and lack of transparency, making potential deployment difficult. To address these concerns, a federated, explainability-driven FER framework designed to provide trustworthy and privacy-preserving emotion recognition with potential applications in mental health monitoring is proposed in this paper. The proposed lightweight Convolutional Neural Network (CNN) enables real-time inference while preserving high accuracy. Comprehensive evaluations on RAF-DB, ExpW, and FER2013 datasets, show that the proposed model demonstrates improved cross-dataset generalization compared to related works, achieving average accuracies of 75.5% and 74.3% in centralized and federated settings, respectively. Quantitative perturbation-based metrics, including Insertion and Deletion Area Under Curve (IAUC and DAUC), Average Drop (AD), Increase in Confidence (IC), Average Drop in Accuracy (ADA), and Active Pixel Ratio, were employed to objectively evaluate the quality and reliability of the model Grad-CAM++ explanations. The results confirm that model explainability enhances transparency and is directly associated with improved model performance.