Rufeng Guo, Rong Gui, Jun Hu, Pinjun Tang, LZ Cao, Jinghui Zhang, Qiao Jiang
Object detection in high-resolution remote sensing images under complex industrial environments is fundamentally constrained by the inherent limitations of single-modality sensors. Optical imagery is prone to background confusion and pseudo-target interference, while synthetic aperture radar (SAR) imagery suffers from speckle noise and structural ambiguity. This work investigates a critical evaluation gap in multimodal fusion, where traditional image-level quality metrics do not consistently reflect downstream detection performance. To address this issue, we propose a task-oriented framework termed the Multi-Source Fusion for Enhanced Object Detection Network (MSFE-Net). The proposed method integrates pixel-level optical–SAR fusion with a YOLOv11-based detector, enabling the learning of task-relevant representations by exploiting complementary optical spectral cues and SAR scattering characteristics. Extensive experiments are conducted across multiple fusion strategies and representative detection architectures on two industrial datasets covering oil tanks and photovoltaic arrays. The results consistently reveal a nonlinear decoupling between image-level fusion metrics and detection accuracy, indicating that improvements in global statistical image quality do not necessarily lead to superior task performance. Furthermore, the proposed framework demonstrates improved robustness in complex scenarios involving multi-scale and weak targets. Specifically, MSFE-Net achieves 99.1% mAP@50 for oil tank detection (19.5% improvement over SAR-only baselines) and 90.2% mAP@50 for photovoltaic array detection, with stable performance across different evaluation settings. These results highlight the importance of task-oriented evaluation in multimodal remote sensing fusion and suggest that downstream detection performance provides a more reliable criterion than conventional image-quality metrics.