科研速览继续刷下去 →
◆ IEEE Transactions on Intelligent Transportation Systems2025-11-07· Computer science

Causality-Driven Explainable Multimodal Fusion With Visual-Text Parallel Computing for Cloth-Changing Pedestrian Re-Identification

Xiejing Yin, Hua Han, Kaiyu Xu, Li Huang, A. A. M. Muzahid

原始摘要(原文)
Pedestrian trajectory prediction and behavior analysis are crucial in intelligent transportation systems (ITS), and the key issue is to effectively achieve pedestrian re-identification (ReID) from cross-camera networks. But however, clothing changes can greatly affect the accuracy of recognition. Since clothing and identity have a complex relationship, existing models inadequately separate identity features from clothing-induced bias through causal analysis. Moreover, current multi-modality methods often overlook the descriptive attributes of individuals in the original RGB images. Therefore, this paper proposes a Causal Textual Visual Network (CTVNet) for Cloth-changing ReID. CTVNet comprises three branches: clothing, identity, and a parallel textual branch, with the textual branch being the primary contribution. This branch introduces two novel modules including an Attribute Extraction and Masking (AEM) module and a Multi-modality Coordination Network (MCN) to unify attribute descriptions with the original RGB images. The parallel textual branch extracts text descriptions unrelated to clothing, fuses them with visual features, and subtracts from identity features to eliminate redundant clothing information, enabling causal intervention. The AEM module masks color and clothing information in descriptive attributes, while the MCN integrates features across tokens and channels, leveraging their effectiveness. For the first time, causal intervention is combined with textual attributes in multi-modality cloth-changing ReID (CC-ReID). Furthermore, masked attribute descriptions are combined with visual features fused by the MCN to reduce the influence of clothing and eliminate residual clothing bias through causal intervention. Extensive experiments on two standard CC-ReID datasets demonstrate the superiority of the proposed CTVNet. Both qualitative and quantitative results are provided. Additionally, several ablation studies are conducted on each component of the proposed model to demonstrate their effectiveness.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Causality-Driven Explainable Multimodal Fusion With Visual-Text Parallel Computing for Cloth-Changing Pedestrian Re-Identification — 科研速览 Science Skim