科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Transformative Technologies and Sustainable Development2025-10-06· Computer science

Fusion of Vision Transformer and Convolutional Neural Network for Explainable and Efficient Histopathological Image Classification in Cyber-Physical Healthcare Systems

Mohammad Ishtiaque Rahman

原始摘要(英文原文)· Original abstract
Abstract Accurate and interpretable classification of breast cancer histopathology images is critical for early diagnosis and treatment planning. This study proposes a hybrid deep learning model that integrates convolutional neural networks (CNNs) with a Vision Transformer (ViT) to jointly capture local texture patterns and global contextual features. The fusion architecture is evaluated on two publicly available datasets: BreakHis and the invasive ductal carcinoma (IDC) dataset. Results demonstrate that the ViT+CNN model consistently outperforms standalone CNN and ViT models, achieving state-of-the-art accuracy while maintaining robustness across datasets. To assess the feasibility of deployment in real-world clinical scenarios, we benchmark inference latency and memory usage under both standard and edge-constrained environments. Although the fusion model has higher computational cost, its latency remains within acceptable thresholds for real-time diagnostic workflows. Furthermore, we enhance interpretability by combining Grad-CAM with attention rollout, allowing for transparent visual explanation of the model’s decisions. The findings support the clinical potential of hybrid transformer-convolutional models for scalable, reliable, and explainable medical image analysis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Fusion of Vision Transformer and Convolutional Neural Network for Explainable and Efficient Histopathological Image Classification in Cyber-Physical Healthcare Systems — 科研速览 Science Skim