Aiman Solyman, Ahmed Elazab, Mohamed Rahouti, Ali Alfatemi, Zeinab Mahmoud, Marco Zappatore
Medical image segmentation is crucial for accurate clinical diagnosis and treatment planning. However, it remains highly challenging due to noise, low contrast, and ambiguous boundaries across different imaging modalities. Existing approaches, including CNN-based models such as U-Net and hybrid Transformer architectures, attempt to mitigate these issues through attention mechanisms, multi-scale feature extraction, or post-hoc boundary refinement. Nevertheless, they often exhibit high computational costs or limited generalization when applied to diverse acquisition protocols and clinical scenarios. To address these challenges, we propose MS-GBANet, a hybrid Transformer-Graph-based architecture that integrates a Pyramid Vision Transformer (PVT) encoder for multi-scale feature extraction, Graph Convolution for long-range spatial reasoning, a Boundary-Aware Feature Refinement (BAFR) module for precise edge localization, and a multi-scale decoding strategy with deep supervision to enhance robustness. The proposed model was evaluated on four datasets: Synapse Multi-organ (CT), ACDC (MRI), ISIC 2018 (dermoscopy), and polyp datasets (colonoscopy). MS-GBANet achieved state-of-the-art performance, with average Dice scores of 84.01, 92.10, 91.67, and up to 94.71 on Synapse, ACDC, ISIC 2018, and CVC-ClinicDB, respectively. Additionally, it achieved the lowest HD95 scores among all compared baseline and state-of-the-art methods across various imaging modalities. The code and data are available at: https://github.com/aimanmutasem/MS-GBANet .