科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ medRxiv2026-08-14· radiology and imaging

VPNet: A Unified Explainable Classification Framework Using Probabilistic Circuits and Foundation Models

T. Kusumoto

原始摘要(英文原文)· Original abstract
Deep learning-based classification models can learn and extract complex representations or relationships between classes. If the decision process of such AI models can be visualized in a way that humans can understand, this may create the potential for AI tools to assist in new scientific discoveries and improve trust in AI models. However, visualizations produced by current post-hoc XAI methods, including Grad-CAM and attention maps, remain theoretically questionable. To overcome this limitation, we propose a novel theoretically supported explainable AI model for classification tasks that can handle multimodal information. Unlike conventional black-box classifiers based solely on convolutional neural networks or transformers, our approach combines probabilistic circuits with two vision transformers, DINO and CLIP, enabling probabilistic interpretability at the patch-associated embedding level, encoder level, and modality level. We demonstrate that our proposed method can pass the randomized test for saliency checks, a test that widely used post-hoc XAI methods often fail. Additionally, this study demonstrates strong generalization performance and representational capability across various classification tasks while preserving the model's explainable structure. To overcome this limitation, we propose a novel theoretically explainable AI model for classification tasks that can handle multimodal and multidimensional information. Unlike traditional classification models based on convolutional neural networks or transformers, our approach combines probabilistic circuits with two vision transformers, DINO and CLIP, enabling probabilistic interpretability at the patch level, encoder level, and modality level. We demonstrate that our proposed method can overcome the randomized test for saliency checks, a test that current state-of-the-art XAI models often fail. Additionally, this study shows the explainability of the model's decisions at the patch, encoder, and modality levels, while achieving strong generalization performance across various classification tasks.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

VPNet: A Unified Explainable Classification Framework Using Probabilistic Circuits and Foundation Models — 科研速览 Science Skim