科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural networks : the official journal of the International Neural Network Society2026-08-31

Backdooring rationalization plus.

Lei Wu, Lingxiao Kong, Fan He, Bo Wang, Yaofeng Su, Xiaoshuang Wang, Xuping Jiang

原始摘要(英文原文)· Original abstract
Rationalization models have recently garnered significant attention for enhancing the interpretability of natural language processing by first using a generator to select the most relevant pieces from the text with respect to the label, before passing the text input to the predictor. However, the robustness of the rationalization models is not sufficiently investigated. Specifically, this paper explores the robustness of rationalization models against backdoor attacks, which has been ignored by previous studies. Surprisingly, we find that conventional backdoor attack techniques fail to inject triggers into the rationalization model because its generator can filter out bad triggers. Considering this, we further propose a novel backdoor attack method named as BadRNL designed specially for the rationalization models. The core idea of BadRNL is first to search for the personalized trigger for each specific dataset and then manipulate the rationales and labels to conduct attacks. Besides, BadRNL controls the order of sample learning through poison-priority sampling strategies. Experimental results across five diverse NLP datasets-FEVER, MultiRC, Beer, Hotel, and Movie-demonstrate that BadRNL consistently achieves near 100% attack success rates (ASR). Crucially, the method maintains high classification accuracy and rationale quality on benign samples, highlighting a significant security risk for interpretable NLP systems deployed in real-world supply-chain scenarios.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Backdooring rationalization plus. — 科研速览 Science Skim