Maryna Mamuta, Ihor Kravchenko, Oleksandr Mamuta
Object detection plays a crucial role in computer vision and many areas of modern life. An outstanding model for this task today is YOLO. However, for specific classes of objects it needs to be trained to distinguish them. For this purpose, transfer learning is used. One promising technique in transfer learning is layer freezing. However, this topic is underexplored, especially for the latest architectures of the YOLO series such as YOLOv11 and YOLO26 despite significant improvements in their architecture, namely replacement of C2f block with C3k2 and adding C2PSA block after SPPF. The question of object detection improvement with YOLOv11 and YOLO26 via layer freezing during transfer learning was addressed in the article. Experiments were conducted on two datasets: medium and small, which were automatically downloaded from the Roboflow API. The layer freezing strategy aimed to check detection improvement with freezing such key backbone layers as C3k2, C2PSA and SPPF and some layers of the neck in comparison with transfer learning without layer freezing in terms of such metrics as precision, recall, mAP@50 and mAP@50-95. All experiments were done in Google Colaboratory Pro environment with NVIDIA A100 GPU (40GB). Experiments were conducted for 50 epochs with an early stopping mechanism (patience 20), batch size 16, learning rate 0.01, optimizer MuSGD for YOLO26 and SGD for YOLOv11 and optimizer AdamW for both models. Results revealed that there is no optimal strategy, but rather empirically driven recommendations that are given in the article. For example, an effective strategy is to freeze layers of the backbone and stop freezing on C3k2 layer for both models with all optimizers. Some benefits are demonstrated by freezing the C2PSA for YOLO26. In the case of a small dataset, it is beneficial to stop freezing at the SPPF block namely for YOLO26 with optimizer MuSGD. Freezing higher layers of the neck did not show significant improvement, however, in some cases, it was beneficial. Results reveal that YOLOv11 demonstrates lower training time and per image inference latency and higher results in precision, recall, mAP@50 and mAP@50-95 metrics than YOLO26.