Seongchan Park, 김진빈, Seunghyun Lee, Soonchul Kwon
Medical artificial intelligence is advancing rapidly and contributing significantly to the early detection and accurate diagnosis of diseases. Although medical image classification is crucial in this diagnostic process, its application in real-world clinical environments demands high accuracy and computational efficiency. To address these requirements, this paper proposes EfficientMedFormer, a novel lightweight hybrid architecture that integrates an anisotropic convolutional block based on the directional decomposition principle with a gated channel-wise anisotropic attention module within a progressive hierarchical structure, effectively fusing local and global features. In the early stages of the model, the anisotropic convolutional block decomposes standard convolutions into directional one-dimensional convolutions. This approach significantly reduces computational cost while maintaining a wide receptive field. Subsequently, in the later stages, the gated channel-wise anisotropic attention module enhances conventional channel attention by independently extracting and directly gating contextual information along the height and width directions. This mechanism dynamically reweights channel importance, effectively preserving spatial information and modeling global relationships. The proposed model achieves comparable or superior performance compared with widely used models in the medical domain across classification tasks on the MedMNIST v2, PAD-UFES-20, Kvasir, and Brain Tumor MRI datasets, achieving a maximum accuracy of 99%. Notably, EfficientMedFormer attains this top-tier performance with only 2.39M parameters and 0.28 GFLOPs, while demonstrating an inference latency of 74.6 ms on an Orange Pi 5B edge device These quantitative results underscore the high practicality and generalizability of the proposed model, thereby advancing the clinical applicability of medical artificial intelligence systems. • A novel lightweight hybrid architecture for medical image classification. • Directional decomposition of operations minimizes computational costs. • Achieves high performance with only 2.39M parameters and 0.28 GFLOPs. • Shows SOTA performance on MedMNIST, Kvasir, PAD-UFES-20 and Brain Tumor MRI datasets. • Demonstrates top-tier inference speeds on various embedded devices.