Sara A Alsalamah, Albatoul Althenayan, Mariam Almutairi, Shada AlSalamah, Hessah A Alsalamah, Thamer Nouh, Bassam Mahboub, Laila Salameh, Moteeb Al Moteri, Chang-Tien Lu
Current research showed significant methodological flaws. AI models based on ensemble architecture and multimodal approaches warrant further exploration. Most studies carry a high risk of bias and rarely undergo external validation, which limits clinical translation. Before clinical implementation, these models still require strict methodological evaluation, prospective multicenter testing, and proper calibration.
INTRODUCTION: Virtual hospital systems support remote triage and diagnosis across distributed healthcare environments by integrating radiographic imaging with structured clinical information for clinical decision making. However, many existing diagnostic frameworks do not fully address real world challenges commonly encountered in scalable virtual healthcare settings. These challenges include severe class imbalance, incomplete clinical attributes, heterogeneous multimodal inputs, and variability in data quality across institutions.
METHODS: This study introduces VHealth-MFusion, a fusion-based hierarchical multimodal deep learning framework that integrates chest X-ray (CXR) imaging and structured clinical data within a unified convolutional neural network (CNN) and multilayer perceptron (MLP) architecture. Using pneumonia Tele-Diagnosis as a use case, the proposed framework combines radiographic feature extraction through a CNN branch with structured clinical representation learning through an MLP branch, followed by late multimodal fusion for multiclass respiratory disease classification. The framework was evaluated using a hierarchical multimodal dataset containing CXR images across multiple diagnostic categories, including Normal, Bacterial, Adenovirus, Influenza, MERS-CoV, and SARS-CoV-2, together with 632 structured clinical variables comprising demographic information, vital signs, symptoms, and laboratory findings.
RESULTS: Under a controlled dataset-consistent comparison against an established hierarchical multimodal CNN baseline, VHealth-MFusion achieved 97.2% classification accuracy, outperforming the baseline accuracy of 95.9%.
DISCUSSION: The findings demonstrate that multimodal integration of radiographic imaging and structured clinical information can improve diagnostic robustness and support more reliable Tele-Diagnosis workflows within virtual hospital environments. Overall, the proposed framework contributes a practically oriented multimodal diagnostic architecture designed to facilitate scalable remote clinical services and informed decision making within digitally enabled care systems where imaging and structured clinical information are jointly available.