科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Remote Sensing2025-12-19· Computer science

SatViT-Seg: A Transformer-Only Lightweight Semantic Segmentation Model for Real-Time Land Cover Mapping of High-Resolution Remote Sensing Imagery on Satellites

D. Shu, Zhan Zhang, Fang Wan, Ru Wang, Bingnan Yang, Yan Zhang, Jianzhong Lu, Xiaoling Chen

原始摘要(英文原文)· Original abstract
The demand for real-time land cover mapping from high-resolution remote sensing (HR-RS) imagery motivates lightweight segmentation models running directly on satellites. By processing on-board and transmitting only fine-grained semantic products instead of massive raw imagery, these models provide timely support for disaster response, environmental monitoring, and precision agriculture. Many recent methods combine convolutional neural networks (CNNs) with Transformers to balance local and global feature modeling, with convolutions as explicit information aggregation modules. Such heterogeneous hybrids may be unnecessary for lightweight models if similar aggregation can be achieved homogeneously, and operator inconsistency complicates optimization and hinders deployment on resource-constrained satellites. Meanwhile, lightweight Transformer components in these architectures often adopt aggressive channel compression and shallow contextual interaction to meet compute budgets, impairing boundary delineation and recognition of small or rare classes. To address this, we propose SatViT-Seg, a lightweight semantic segmentation model with a pure Vision Transformer (ViT) backbone. Unlike CNN-Transformer hybrids, SatViT-Seg adopts a homogeneous two-module design: a Local-Global Aggregation and Distribution (LGAD) module that uses window self-attention for local modeling and dynamically pooled global tokens with linear attention for long-range interaction, and a Bi-dimensional Attentive Feed-Forward Network (FFN) that enhances representation learning by modulating channel and spatial attention. This unified design overcomes common lightweight ViT issues such as channel compression and weak spatial correlation modeling. SatViT-Seg is implemented and evaluated in LuoJiaNET and PyTorch; comparative experiments with existing methods are run in PyTorch with unified training and data preprocessing for fairness, while the LuoJiaNET implementation highlights deployment-oriented efficiency on a graph-compiled runtime. Compared with the strongest baseline, SatViT-Seg improves mIoU by up to 1.81% while maintaining the lowest FLOPs among all methods. These results indicate that homogeneous Transformers offer strong potential for resource-constrained, on-board real-time land cover mapping in satellite missions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

SatViT-Seg: A Transformer-Only Lightweight Semantic Segmentation Model for Real-Time Land Cover Mapping of High-Resolution Remote Sensing Imagery on Satellites — 科研速览 Science Skim