科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Visual Computer2026-02-01· Computer science

Enhancing RGB-IR object detection: a frozen backbone approach with multi-receptive field attention

Bingyu Lu, H. Liu, Hiroshi Watanabe

原始摘要(英文原文)· Original abstract
Recent advancements in multimodal object detection have predominantly relied on end-to-end training paradigms, which, while effective, demand substantial computational resources and risk feature degradation. To address these challenges, we propose a frozen backbone paradigm, preserving pretrained representations as stable semantic anchors for efficient multimodal fusion. Our approach introduces a lightweight multi-receptive field attention (MRFA) mechanism, enhancing feature interaction and representation diversity without exhaustive retraining. Experiments on the FLIR Aligned and M $$^3$$ FD dataset demonstrate consistent improvements over state-of-the-art end-to-end models, highlighting the potential of pretrained backbones coupled with adaptive attention mechanisms for robust multimodal object detection. The project code is released at https://github.com/LuBingyu11/MRFA .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Enhancing RGB-IR object detection: a frozen backbone approach with multi-receptive field attention — 科研速览 Science Skim