Ali Hassan, Johan Johansson, Stefan Schulte, Tingting Zhang, Karen Egiazarian, Mårten Sjöström
Semantic segmentation on resource-constrained hardware remains a key challenge in deep learning, particularly for deployment on edge devices and embedded systems. In this study, we propose FireSegUNet, a lightweight and computationally efficient deep-learning architecture tailored for such environments. The model integrates an optimized inverted bottleneck layer for feature extraction within an encoder–decoder framework, reducing computational complexity by up to 51%. It also improves segmentation accuracy through an efficient squeeze-and-excitation block, while reducing inference time by up to 7.3 × and energy consumption by up to 4.6 × compared to conventional attention mechanisms. Extensive evaluation on diverse fire segmentation datasets demonstrates that FireSegUNet achieves competitive segmentation accuracy while reducing the number of parameters and storage requirements by up to 81%. Additionally, we provide a detailed analysis of the relationship between model complexity metrics and actual inference time, memory usage, and energy consumption. This comprehensive evaluation confirms that FireSegUNet delivers better performance on edge devices and generalizes well to unseen datasets. These findings position FireSegUNet as a practical solution for efficient image segmentation in resource-constrained environments. Although primarily validated on fire segmentation, the modular design of FireSegUNet makes it adaptable to other computer vision tasks. The source code of FireSegUNet will be publicly available at https://github.com/Realistic3D-MIUN/FireSegUNet .