Ouail Choukhairi, Mouad Choukhairi, Ali Choukri, Youssef Fakhri
Semantic segmentation of architectural elements in cultural heritage sites lies at the intersection of computer vision and digital preservation. Moroccan historical monuments spanning mosques, madrasas, royal gates (babs), and mausoleums across Fez, Rabat, Marrakech, Meknes, and Tetouan present unique challenges, including extreme texture ambiguity between weathered wall surfaces and background, pronounced multi-scale variation from individual window and door openings to full rooftop surfaces spanning tens of metres, and severe class imbalance. We introduce MonuSegFormer, a heritage-specific hybrid architecture coupling a pretrained Swin-B encoder (ImageNet-22K) with a Multi-Scale Atrous Fusion (MSAF) module and a CBAM-augmented progressive decoder. Evaluated on the Moroccan Monuments Dataset (MMD, 2686 annotated RGB images, 5 semantic classes, 5 cities) on the held-out test set (403 images), MonuSegFormer achieves mIoU = 81.7%, outperforming SegFormer-B5 (76.4%), Mask2Former (78.9%), and DeepLabV3+ (71.3%). A composite CE+Dice+Focal loss with class-balanced weights addresses severe class imbalance, yielding the largest gains on the minority classes Door (+5.3 p.p.) and Window (+5.3 p.p.) over the next-best Mask2Former. All qualitative results are produced by real trained-model inference on held-out test images.