Abeer S. Almogren, Ghada Moh. Samir Elhessewi, Mukhtar Ghaleb, Hany Mahgoub, Asma A. Alhashmi, Umkalthoom Alzubaidi, Asmaa Mansour Alghamdi, Nojood O. Aljehane
Social Internet of Things (IoT) creates an integrated ecosystem for mental health analysis through video emotion analysis. These systems can be transformed into real-time monitoring tools and support mental health care systems such as online education, interviews, and medical treatments. This research proposed video-based emotion classification with a novel optimized hybrid deep learning model. From the video, facial expression features such as eye movements and lip movements are captured for emotion analysis. Videos are converted into frames and are trained and tested using various baseline deep learning models such as convolutional neural network (CNN), CNN for long short-term memory (LSTM), 3D convolutional neural network (3D CNN), and CNN-LSTM model. We proposed a novel optimizer-based hybrid deep learning model for improving the accuracy of emotion detection—the CNN-LSTM-GWO and CNN-LSTM-PSO. In the final task output, we employ a multitask learning mechanism, setting discrete emotion recognition as the primary task and emotion valence recognition, emotion arousal recognition, and previous information extraction as auxiliary tasks to facilitate practical information sharing across different tasks. Experimental results on the established driver emotion dataset demonstrate that our proposed method significantly improves driver emotion recognition performance, achieving an accuracy of 86.98% and an F1 score of 85.83% in the primary task. This validates the effectiveness of the proposed approach.