Joseph Ngwa, Elie Fute Tagne, Nde Nguti
Accurate building segmentation from high-resolution aerial imagery remains challenging due to variations in building size, shape, background clutter, and class imbalance.This study proposes an Enhanced U-Net architecture for automatic building extraction using the Massachusetts Buildings Dataset and the Inria Aerial Image Labeling Dataset.The proposed model extends the conventional U-Net by incorporating a deeper encoder with Batch Normalization, Attention Gates (AGs) in decoder skip connections, and an Atrous Spatial Pyramid Pooling (ASPP) module for multi-scale contextual feature extraction.Different ASPP dilation-rate configurations were investigated, with dilation rates of 3, 6, and 9 yielding the best performance.To handle the high spatial resolution of the imagery, a patch-based training strategy with 50% overlap was adopted, using patch sizes of 256 × 256 and 512 × 512 for the Massachusetts and Inria datasets, respectively.The model was optimized using a hybrid Binary Cross-Entropy (BCE) and Dice loss function to balance pixel-level classification and region-overlap accuracy.Experimental results demonstrate that the proposed Enhanced U-Net outperformed U-Net, DeepLabv3+, and HRNet.On the Massachusetts Buildings Dataset, it achieved an IoU of 0.737 and an F1-score of 0.848.On the Inria dataset, it achieved an IoU of 0.799 and an F1score of 0.888.These results demonstrate the effectiveness and generalization capability of the proposed architecture for aerial building segmentation.