科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Nature biomedical engineering2026-08-17

An explainable biomedical foundation model via large-scale concept-enhanced vision-language pretraining.

Yuxiang Nie, Sunan He, Yequan Bie, Yihui Wang, Zhixuan Chen, Shu Yang, Zhiyuan Cai, Linshan Wu, Hongmei Wang, Xi Wang, Ngai Shing Cheng, Luyang Luo, Mingxiang Wu, Haibo Jin, Xian Wu, Ronald Cheong Kin Chan, Yuk Ming Lau, Zhengyu Zhang, Sushan Xiao, Can Yang, Yinghua Zhao, Xiaohui Duan, Li Zhang, Li Liang, Yefeng Zheng, Pranav Rajpurkar, Hao Chen

原始摘要(英文原文)· Original abstract
Artificial intelligence for medical imaging is required to be accurate and interpretable to clinicians. However, current multimodal biomedical foundation models often prioritize performance over explainability. Here we present ConceptCLIP, an explainable biomedical foundation model that achieves state-of-the-art diagnostic accuracy while delivering human-interpretable explanations across diverse imaging modalities. We curate MedConcept-23M, a large-scale dataset comprising 23 million biomedical image-text-concept triplets. Leveraging this dataset, we pretrain ConceptCLIP via joint image-text and region-concept alignment for precise and interpretable medical image analysis. Across a large-scale benchmark covering 78 datasets in 10 imaging modalities, ConceptCLIP demonstrates superior diagnostic performance while providing human-understandable explanations. In a clinician user study spanning three modalities, the concept-based explanations provided by ConceptCLIP help clinicians verify model predictions and identify potential errors. As an explainable biomedical foundation model, ConceptCLIP represents a critical milestone towards the widespread clinical adoption of AI, thereby advancing trustworthy AI in medicine.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

An explainable biomedical foundation model via large-scale concept-enhanced vision-language pretraining. — 科研速览 Science Skim