科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Affective Computing2026-01-01· Sarcasm

InterARM: Interpretable Affective Reasoning Model for Multimodal Sarcasm Detection

Yue Tan, Rui Mao, Xuzhao Shi, E. Cambria

原始摘要(英文原文)· Original abstract
Multimodal sarcasm detection (MSD) aims to identify sarcastic expressions by integrating multimodal and contextual cues to capture cross-modal semantic inconsistencies. However, existing studies face challenges: fine-grained sarcastic cues are implicit and dispersed, large generative models suffer from gradient vanishing in classification tasks, and cross-domain generalization remains limited. To address these limitations, we propose InterARM, an interpretable affective reasoning model that introduces a structured sarcasm reasoning paradigm. Specifically, we introduce a three-stage training strategy based on curriculum learning, consisting of 1) sarcasm classification learning, 2) structured sarcasm reasoning learning, and 3) step-selective hierarchical reward reinforcement learning. The model generates reasoning chains across modalities to analyze sentiment polarity and semantic conflicts. We also construct the MSD-CoT dataset with 3,200 image-text pairs with 19,200 human-annotated reasoning steps. Experiments show that InterARM achieves superior performance on the in-domain and out-of-domain datasets, outperforming larger MLLMs such as Qwen2.5-VL-7B and GPT-5-mini while maintaining high interpretability and strong cross-domain generalization.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

InterARM: Interpretable Affective Reasoning Model for Multimodal Sarcasm Detection — 科研速览 Science Skim