Bakhita Salman, Eithar Yassin, Deepak Ganta, Hermes Luna
Despite advances in deep learning, brain tumor detection from MRI continues to face major challenges, including the limited robustness of single-modality models, the computational burden of transformer-based architectures, opaque fusion strategies, and the lack of efficient binary screening tools. To address these issues, we propose a lightweight multimodal CNN framework that integrates T1, T2, and FLAIR MRI sequences using modality-specific encoders and a channel-wise fusion module (concatenation followed by a 1 × 1 convolution). The pipeline incorporates U-Net-based segmentation for tumor-focused patch extraction, improving localization and reducing irrelevant background. Evaluated on the BraTS 2020 dataset (7500 slices; 70/15/15 patient-level split), the proposed model achieves 93.8% accuracy, 94.1% F1-score, and 19 ms inference time. It outperforms all single-modality ablations by up to 5% and achieves competitive or superior performance to transformer-based baselines while using over 98% fewer parameters. Grad-CAM and LIME visualizations further confirm clinically meaningful tumor-region activation. Overall, this efficient and interpretable multimodal framework advances scalable brain tumor screening and supports integration into real-time clinical workflows.