Yiming Li, Mingyan Zhu, Shu-Tao Xia, Zhifeng Li, Zhan Qin, Dacheng Tao
Backdoor attacks intend to inject hidden backdoor into the deep neural networks (DNNs), such that the predictions of infected models will be maliciously changed if the hidden backdoor is activated by the attacker-specified trigger pattern. Since the infected models behave normally on predicting benign samples, the backdoor attack is stealthy and therefore a serious threat to practical applications of DNNs. Currently, most existing backdoor attacks adopted the setting of static trigger, i.e ., triggers across the training and testing images follow the same appearance and are located in the same area. In this paper, we revisit this attack paradigm by analyzing trigger characteristics. We demonstrate that this attack paradigm is vulnerable when the trigger in testing images is not consistent with the one used for training. As such, those attacks are far less effective in the physical world, where the location and appearance of the trigger contained in digitized test samples may be different from that of the one used for training. Besides, we introduce a plug-in attack enhancement module during training, inspired by the expectation over transformation (EOT), to alleviate such inconsistency vulnerability. Based on this plug-in module, we also reveal that the widely adopted data augmentation may exacerbate the security risks of backdoor attacks, although it can enhance model performance. Moreover, we evaluate our methods on multiple benchmark datasets to verify their effectiveness. We hope that our work could inspire more explorations on the properties of backdoor attacks, to facilitate the design of more robust and secure DNNs.