Binghong Li, Chuxin Xiao
Link prediction in multimodal knowledge networks jointly exploits topology, text, and images, but visual backgrounds may introduce spurious correlations that reduce robustness under distribution shift. We propose MKGC-CSR, a causal deconfounding framework that models visual context as a confounder and approximates backdoor adjustment through a visual concept dictionary and stratified causal attention. The dictionary discretizes the continuous visual context space into prototypes, while the attention module aggregates prototype-conditioned representations using empirical priors. Experiments on FB15k-237-IMG, WN18-IMG, DB15K, and MKG-W show that MKGC-CSR achieves competitive clean-data performance, including an MRR of 0.405 on FB15k-237-IMG. Its margins over the strongest recent baselines are modest (0.003-0.006 MRR) but statistically significant across five random seeds, with limited overhead (+ 2.3 ms inference latency and + 3M parameters). The main benefit lies in robustness: under Gaussian noise, occlusion, background replacement, illumination change, and style shift, MKGC-CSR exhibits consistently smaller performance degradation than clean-trained baselines, and remains more robust than noise-augmented baselines under Gaussian corruption. Sensitivity analysis further shows that the gains are mainly attributable to background-dominated prototypes, supporting the proposed visual-confounder assumption within its stated limits.