Ayokunle Olalekan Ige, Daniel Ayo Oladele, Malusi Sibiya
Human Activity Recognition (HAR) using wearable sensor data plays a vital role in health monitoring, context-aware computing, and smart environments. Many existing deep learning models for HAR incorporate MaxPooling layers after convolutional operations to reduce dimensionality and computational load. While this approach is effective in image-based tasks, it is less suitable for the sensor signals used in HAR. MaxPooling introduces a form of temporal downsampling that can discard subtle yet crucial temporal information. Also, traditional CNNs often struggle to capture long-range dependencies within each window due to their limited receptive fields, and they lack effective mechanisms to aggregate information across multiple windows without stacking multiple layers, which increases computational cost. In this study, we introduce Retentive-HAR, a model designed to enhance feature learning by capturing dependencies both within and across sliding windows. The proposed model intentionally omits the MaxPooling layer, thereby preserving the full temporal resolution throughout the network. The model begins with parallel dilated convolutions, which capture long-range dependencies within each window. Feature outputs from these convolutional layers are then concatenated along the feature dimension and transposed, allowing the Retentive Module to analyze dependencies across both window and feature dimensions. Additional 1D-CNN layers are then applied to the transposed feature maps to capture complex interactions across concatenated window representations before including Bi-LSTM layers. Experiments on PAMAP2, HAPT, and WISDM datasets achieve a performance of 96.40%, 94.70%, and 96.16%, respectively, which outperforms the existing methods with minimal computational cost.