科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural networks : the official journal of the International Neural Network Society2026-09-23

Context-aware causal reasoning for explanatory visual question answering.

Jiali Miao, Xiaoling Huang, Kui Yu, Yecheng Shi, Xiang Wang, Fuyuan Cao

原始摘要(英文原文)· Original abstract
Explanatory Visual Question Answering (EVQA) is a multimodal reasoning task that answers visually relevant natural language questions and generates user-friendly multimodal explanations. Although current EVQA methods can generate grammatically sound and vocabulary appropriate explanations, they still face the challenge of explanatory illusion (i.e., wrong explanations arrive at correct answers). To address this issue, we propose a novel Context-Aware Causal Reasoning (CACR) algorithm. Specifically, we first leverage a Multimodal Large Language Model and design prompting strategies to generate high-quality image descriptions as complementary evidence for helping generate correct explanations. Then, we propose a Structural Causal Model (SCM) to establish the causal relationship among the description, the explanation, and the answer. Finally, we transform the SCM into a deep variational inference network framework that enables causal reasoning to ensure that the explanations are grounded in the visual evidence related to the answer. Extensive experiments show that CACR significantly reduces the risk of explanatory illusion compared to six state-of-the-art methods, outperforming the best baseline by 2.22% on average in terms of explanation quality.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Context-aware causal reasoning for explanatory visual question answering. — 科研速览 Science Skim