Yinuo Wang, Kai Chen, Yue Zeng, Meng Cai, Chao Pan, Zhouping Tang
Comparison of responses from six multi-modal large language models on a random case, which illustrates a fracture in the patient’s left skull due to trauma, with subdural hemorrhage (SDH) evident in the left hemisphere and intraparenchymal hemorrhage (IPH) in the right hemisphere. • Multi-modal large language models (MLLMs) show potential in intracranial hemorrhage (ICH) diagnosis and treatment with visual and interactive capabilities. • Proprietary and open-source MLLMs still underperform in ICH binary/subtype classification compared to trained deep networks. • MLLMs require fine-tuning for improved ICH subtyping and related tasks. Objective: Accurate identification of intracranial hemorrhage (ICH) subtypes on non-contrast CT is crucial for prognosis and treatment but remains challenging due to low contrast and blurred boundaries. This study evaluates the zero-shot performance of multi-modal large language models (MLLMs) versus traditional deep learning in ICH detection and subtyping. Methods: Using 192 NCCT volumes from the RSNA dataset, we compared MLLMs (GPT-4o, Gemini 2.0 Flash, Claude 3.5 Sonnet V2) with deep learning models (ResNet50, Vision Transformer). MLLMs were prompted for ICH presence, subtype, localization, and volume estimation. Results: Traditional deep learning models outperformed MLLMs in both ICH detection and subtyping. For subtyping, MLLMs showed lower accuracy, with Gemini 2.0 Flash achieving a macro-averaged precision of 0.41 and F1 score of 0.31. Conclusions: While MLLMs offer enhanced interpretability through language-based interaction, their accuracy in ICH subtyping remains inferior to deep learning networks. Further optimization is needed to improve their utility in three-dimensional medical imaging.