Félix Michaud, Jérôme Sueur, Frédéric Sèbe, Maxime Le Cesne, Sylvain Haupert
Automatic sound classification is a promising approach to analyze the spatio-temporal dynamics of natural sound sources. Convolutional neural networks (CNNs) have proved to be particularly effective for the classification of animal sounds. However, even if these algorithms are more efficient than their predecessors, CNNs need to show stable performances despite various sources of variability. The signal-to-noise ratio (SNR) has been identified in the literature as a key parameter in the performance of sound classification algorithms. In this study, we investigated the influence of the SNR during the training and inference of a custom-built CNN for the mono-specific classification of Boreal Owl ( Aegolius funereus ) vocalizations. The repetitive, stereotyped nocturnal vocalizations of this species are ideal for isolating the effects of the SNR. Experiments showed that even if the custom-built model was trained with very low SNRs, it still very poorly detected low SNR vocalizations during testing. Performance drops per SNR of the custom-built model were comparable to those of the popular model BirdNET. If the custom-built model outperformed BirdNET overall, both models seemed to have similar performance drops as soon as vocalizations had a SNR inferior to 3 dB. Finally, the custom-built model classified 6 years of audio recordings from 4 autonomous recorders in the Risoux forest (France) and highlighted the variations in the Boreal Owl vocalization activity on a yearly, monthly and hourly scale.