Haiqi Zhang, Hao Tang, Yanpeng Sun, Zechao Li
Vision-Language Pre-training (VLP) models exhibit pronounced vulnerability to multimodal adversarial examples, necessitating rigorous robustness research, particularly for transferable attacks in black-box scenarios. Current research predominantly enhances attack transferability across VLP models by diversifying image and text inputs. However, during adversarial example generation, these methods often prioritize amplifying inter-modal semantic discrepancies (i.e., modality-discrepancy features) while overlooking model-specific semantic features critical to transferable attacks. To address this limitation, we pro pose a transferable Gradient Pruning Interactive Attack (GPI Attack), which integrates gradient-pruned image perturbations with semantic-oriented text perturbations through modality interaction. For image attacks, extreme backpropagated gradients may cause adversarial examples to highlight certain model specific features, leading to poor transferability. To suppress this feature, the textual modality guides the pruning of extreme gradients within intermediate VLP blocks, and these pruned gradients are subsequently employed to direct the generation of adversarial images. For text attacks, we consolidate the perturbation process solely at the embedding level, which reduces semantic discrepancies across hierarchical structures and significantly enhances the generalizability of adversarial texts. Experimental results demonstrate the effectiveness of GPI-Attack in image text retrieval tasks on multimodal datasets such as Flickr30K and MSCOCO. Additionally, the proposed gradient pruning technique is plug-and-play, showing performance improvements even when applied to baseline methods, indicating its potential as a valuable enhancement for attack performance.