Chentong Wang, Jincheng Gao, Fei Zhu, Abderrahim Halimi, Cédric Richard
Transformers have shown significant success in hyperspectral unmixing (HU). However, challenges remain. Transformer-based unmixing networks, built on Vision Transformer (ViT) or Swin Transformer, struggle to effectively capture essential multi-scale and long-range spatial correlations. Additionally, these networks predominantly rely on the linear mixing model, lacking the flexibility to accommodate scenarios with significant nonlinear effects. To address these limitations, we propose a multi-scale Dilated Transformer-based unmixing network for nonlinear HU (DTU-Net). Its encoder integrates two branches: a spatial branch mainly employing Multi-Scale Dilated Attention (MSDA) to uniquely capture intricate multi-scale and long-range spatial correlations via adaptive receptive fields, and a spectral branch utilizing 3D-CNNs with channel attention. This design enables comprehensive extraction as well as integration of multi-level spatial and spectral features. The decoder is specifically designed to accommodate both linear and nonlinear mixing. It explicitly models the polynomial post-nonlinear mixing model (PPNMM) by learning nonlinear coefficients as pixel-wise features, which enhances interpretability by directly reflecting the pixel-level nonlinear mixing strength. Experiments on synthetic, ray tracing, and real datasets validate the effectiveness of the proposed DTU-Net, demonstrating its superior performance compared to both PPNMM-derived and advanced unmixing networks. The code is available at: https: //github.com/ChentongWang/DTU-Net.