Zaoping Zhong, Chengbin Liang, Ming Yang, Deguang Wang
Deep learning models, especially hybrid models combining convolutional neural networks (CNN) and Transformer, introduce new ideas for hyperspectral image (HSI) classification. However, the complex structure and high spectral dimension of HSI data often affects classification accuracy and computational complexity. Therefore, a novel lightweight CNN-Transformer network (LCTNet) model is proposed. LCTNet consists of three core modules designed to improve HSI classification performance while maintaining a lightweight structure. The spectral extraction-dimension reduction module extracts essential spectral features from the hyperspectral data, simultaneously reducing data dimensions to decrease the computational load for subsequent processing. The spatial-spectral feature marking module then enhances crucial spatial and spectral features, facilitating the efficient fusion and utilization of spatial-spectral information. Meanwhile, the dynamic sparse transformer module dynamically adjusts the proportion of attention sparsity, which reduces redundant computation but also preserves critical spectral-spatial global dependencies. These three modules enable LCTNet to effectively improve the classification accuracy and reduce the complexity. Experiments demonstrate that LCTNet achieves overall accuracy of 99.31%±0.26, 99.42%±0.13, 98.13%±0.15, and 98.90%±0.05 on four public datasets, respectively. Compared with nine other methods, LCTNet attains almost all the highest accuracy with relatively low computational complexity. Additionally, the ablation experiment also further verified the function of each module of LCTNet.