Mahdi Samaee, Mehran Yazdi, Daniel Massicotte
Automated sleep stage classification from polysomnography remains challenged by limited modeling of long-range temporal dependencies, suboptimal multimodal EEG–EOG fusion, and insufficient interpretability of deep learning approaches. We propose NeuroLingua, a language-inspired framework that models sleep as a structured physiological sequence. Each 30-second epoch is decomposed into overlapping 3-second subwindows (“tokens”) using a CNN-based tokenizer. Hierarchical temporal dependencies are captured through dual-level Transformers: an intra-segment encoder for local dynamics and an inter-segment encoder integrating seven consecutive epochs (3.5 minutes) to model extended context. Modality-specific embeddings from EEG and EOG channels are fused via a Graph Convolutional Network (GCN), enabling structured multimodal integration. NeuroLingua is evaluated on the Sleep-EDF Expanded and ISRUC-Sleep datasets. On Sleep-EDF, the model achieves 85.3% accuracy, 0.800 macro-F1, and 0.796 Cohen’s κ, demonstrating state-of-the-art performance. On ISRUC, it attains 81.9% accuracy, 0.802 macro-F1, and 0.755 κ, remaining competitive with published baselines across overall and per-class metrics. Attention analysis further highlights physiologically meaningful temporal patterns, supporting improved interpretability. By combining hierarchical sequence modeling with graph-based multimodal fusion, NeuroLingua provides a structured and extensible framework for transparent and clinically informed automated sleep staging.