Hanyu Zhang, Zhongde Zhang, Weiping Liu
Continuous, non-invasive fish monitoring supports aquatic animal management, biodiversity assessment, and sustainable aquaculture, but embedded deployment requires a careful balance among accuracy, speed, memory, and computation under visually degraded underwater conditions. We developed ULFD-YOLO, an ultra-lightweight detector derived from YOLOv11n through coordinated redesign of the backbone, neck, and detection head. The model combines a custom convolutional MobileNetV4-tiny backbone, a hypergraph-based multi-scale fusion neck, and a lightweight MBConv head with channel attention. Experiments were conducted on Fish-BJ, an in-house dataset of 3402 images covering 21 species-informed aquarium-fish detection categories, and on a deliberately difficult 1180-image WildFish subset after dataset-specific training. On Fish-BJ, ULFD-YOLO achieved 0.960 mAP@0.5 and 0.732 mAP@0.5:0.95 with 1.3 M parameters, 2.6 GFLOPs, and a 3.0 MB model file, reducing parameters and computation by 50.0% and 58.7% relative to YOLOv11n. Bootstrap resampling yielded 95% confidence intervals of 0.946-0.973 and 0.638-0.821 for the two metrics, respectively. The model achieved 0.803 mAP@0.5 on WildFish and 19-24 FPS at 448 × 640 on a Jetson Orin Nano under its 15 W nvpmodel power mode. These results establish a practical accuracy-efficiency trade-off for embedded fish monitoring rather than peak localization accuracy.