Lei Zhu, Qifeng Yue, Aiai Huang, Jianhai Zhang, Peng Yuan
Motor imagery (MI)-based brain-computer interfaces (BCIs) decode EEG signals into control commands. However, fine-grained MI decoding within the same limb remains challenging due to highly similar neural patterns. This paper proposes a Contrastive Learning Network based on a Multi-Scale Transformer (CLMT-Net) for fine-grained MI decoding. CLMT-Net integrates multi-scale temporal convolution, FFT-based frequency fusion, spatial convolution, and dual-path Transformer to learn complementary EEG representations. Supervised contrastive learning further improves feature discrimination. On the MI-2 dataset, CLMT-Net achieves an accuracy of 76.13 ± 6.77% with a 95% confidence interval of [73.33, 78.92], demonstrating competitive performance for same-limb MI decoding.