Ying Chen, Chenglong Ma, Qiongqiong Li, Fang Yan, Yirong Chen, Tianbin Li, Jin Ye, Ming Hu, Yuxiang Lin, Yanjun Li, Guoan Wang, Huihui Xu, Hui Dong, Xiang Wang, Xiaoxiao Xu, Yanyan Zhou, Xia Zhu, Sen Yang, Xiyue Wang, Lu Zhang, Yu Qiao, Rongshan Yu, Junjun He, Yuanfeng Ji
Multimodal artificial intelligence, albeit showing great potential in computational pathology, remains limited to isolated patch-level interpretation and often fails to analyze gigapixel-scale whole-slide images (WSIs) essential for clinical utility. Here we present SlideChat, a multimodal generative artificial intelligence assistant for whole-slide computational pathology across cancer types. SlideChat integrates patch-level and slide-level pathology encoders with a pretrained large language model. Using 274,233 multimodal instruction samples, SlideChat is trained to learn the associations between WSIs and diagnostic reports and interpret complex queries in clinical practice. Evaluated on 8,836 closed-ended questions, 129 open-ended questions and 3,149 WSI reports from five cohorts spanning 31 cancer types, SlideChat outperformed leading baselines by 19.1% in closed-ended accuracy and by 7.7% in report-generation Metric for Evaluation of Translation with Explicit Ordering score and received the highest expert ratings across five dimensions in open-ended question answering, showing the potential to enhance diagnostic workflows, medical education and clinical decision-making.