Yi Wang, Jiayong Yan, Chaowen Li, Hui Yang
These results demonstrate that the proposed multi-modal framework provides an effective and scalable solution for sleep staging using cardiorespiratory signals, with good cross-dataset generalization capability and potential for practical wearable applications.
BACKGROUND: Sleep staging is a fundamental task in clinical sleep medicine and wearable health monitoring. Traditional sleep staging methods primarily rely on polysomnography which provides high accuracy but involves complex instrumentation and high cost, limiting its application in home-based monitoring and large-scale screening.
METHODS: To enable a more convenient assessment of sleep structure, this study proposes a multi-modal deep learning approach based on instantaneous heart rate (IHR) and instantaneous breathing rate (IBR). In this study, a TCN-BiGRU-SE network is developed, where the temporal convolutional network (TCN) extracts multi-scale temporal features, the bidirectional gated recurrent unit (BiGRU) captures long-term dependencies, and the squeeze-and-excitation (SE) module enhances feature representation via channel-wise attention. The proposed method is evaluated on two public datasets, Sleep Heart Health Study and MESA, for four-class sleep staging (Wake, rapid eye movement, Light, and Deep).
RESULTS: The proposed method achieves an overall accuracy of 82.19% with a Kappa coefficient of 0.74 on the SHHS dataset, and 80.44% accuracy with a Kappa of 0.70 on the MESA dataset. It outperforms traditional HRV/RRV-based approaches and single-modality models, while maintaining stable performance across datasets.
CONCLUSION: These results demonstrate that the proposed multi-modal framework provides an effective and scalable solution for sleep staging using cardiorespiratory signals, with good cross-dataset generalization capability and potential for practical wearable applications.