Yifan Xi, Ting Lu, Jiacheng Lu, Xudong Kang, Shutao Li
Visible and infrared light images reflect object characteristics in different aspects, which has attracted much attention for object detection in recent years. Nevertheless, the existing multimodal detection networks may fail in the absence of modality. In order to address this problem, a new cross-modal knowledge distillation network (CMKD-net) is proposed for oriented object detection in visible and infrared images. In brief, a teacher-student (T-S) learning network is constructed, where the T-network aims to learn a discriminative feature representation from multimodal images and then guides the S-network training with incomplete modality. Here, multi-dimensional feature distillation (MDFD) and inter-instance relation distillation (IIRD) are designed for cross-modal knowledge propagation. Specifically, the MDFD considers pushing the T-S networks to learn a similar data distribution and feature representation through channel-spatial dimensional feature consistency constraints. The IIRD contributes to retaining the relation structure between individual targets in multimodal images via inter-instance relation modeling and similarity distance measurement. Moreover, to avoid the bias of feature extraction caused by discrete quantization in traditional pooling operations, a rotation-adaptive RoI Pooling (RA-RoI Pooling) is introduced by calculating the continuous double integral within each bin of oriented objects. Ablation experiments and comparison experiments on the VEDAI and DroneVehicle datasets can demonstrate the effectiveness of the proposed CMKD-net.