Nataliya Bilous, Vladyslav Malko, Iryna Ahekian, Marcus Frohme
The integration of artificial intelligence into video-based human behavior analysis enables contactless and continuous monitoring of both motor dynamics and facial reactions. This paper proposes a dual-stream multimodal framework for synchronized modeling of facial expression dynamics and skeletal motion during physical movement from monocular RGB video. The framework consists of two coordinated streams: the motor stream, based on 2D skeletal keypoints, and the facial stream, which extracts features associated with discomfort and affective responses. Person and face detection are performed using YOLO11, while specialized deep learning models handle pose estimation and facial expression recognition. Temporal dependencies and cross-modal relationships are modeled via a bidirectional LSTM, enabling unified temporal modeling of skeletal and facial dynamics. This novel approach allows investigation of how physical movement patterns relate to facial reactions during dynamic activities. By integrating heterogeneous facial and skeletal features in a synchronized temporal model, the framework enables consistent cross-modal analysis of dynamic human behavior. The framework was trained and validated using FER2013 and AffectNet for facial expression recognition, and UI-PRMD and FineRehab for skeletal motion modeling. It achieves 91.2% accuracy in facial expression classification, 94.8% mean Intersection over Union for human detection, and an F1 score of 0.89 for multimodal state assessment. Operating in real-time at 18–28 FPS on standard GPU hardware without requiring wearable sensors, the framework supports applications in behavioral monitoring and safety analysis.