Jintao Sun, Jian Wang, Ming Yu
Automatic classroom student behavior recognition faces three major challenges: the coexistence of small targets and complex backgrounds, the conflict between details and semantics, and the interference from low-quality samples. To address these issues, this paper proposes a scene-driven integrated optimization framework, termed Hyper-YOLO-G. In the backbone network, a Squeeze-and-Excitation (SE) attention module is embedded to purify features and thereby enhance the channels of key behaviors. In the neck, a bidirectional weighted feature pyramid network (BiFPN) is introduced to perform scale-balanced multi-scale fusion. Moreover, the Wise-IoU v3 (WIoU v3) loss function is adopted to calibrate gradients and reduce the negative impact of low-quality samples. Experimental results on a public classroom behavior dataset show that Hyper-YOLO-G achieves 73.7% mAP@50 (mean average precision at IoU threshold 0.5), which is 4.9% higher than the baseline Hyper-YOLO. Ablation studies and generalization experiments on an independent multi-class dataset further demonstrate the combined effectiveness of each module and the generalization ability of the framework. This study provides a reliable technical path for high-precision, lightweight behavior recognition in complex classroom environments.