Bo Cao, Yue Zhang, Qun Gai, Mengze Zhang, Fan Yu, Mengmeng Feng, Senhao Zhang, Jinxu Liu, Feng Shi, Jie Lu
Multimodal learning has gained attention in recent years due to its ability to effectively utilize data features from various modalities. Diagnosing the vulnerability of atherosclerotic plaques directly from carotid 3D MRI images is challenging for both radiologists and conventional 3D vision networks. In clinical practice, radiologists assess patients using a multimodal approach that incorporates various imaging modalities and domain-specific expertise, paving the way for the creation of multimodal diagnostic networks. In this study, we proposed an effective framework to leverage radiologists' domain knowledge to improve the automated diagnosis of carotid plaque vulnerability through variational inference and multimodal knowledge distillation (VMD). This framework excels in harnessing cross-modality prior knowledge from limited image annotations and radiology reports within training data, thereby enhancing the diagnostic network's accuracy for unannotated 3D MRI images. We validated the proposed VMD framework on our in-house dataset, demonstrating its effectiveness.