Nisreen Talib Abdulhusein, Basheera M. Mahammod
Classifying environmental sounds is considered a challenging task because of the complex nature of acoustic signals. This classification involves identifying the acoustic flow associated with different sounds in the environment. Several factors affect it, such as the distances between the source and the destination, voice interference, and the huge number of sound sources. These problems have been addressed in this work, where a hybrid deep learning technique is proposed to improve the classification of sounds, focusing on fundamental audio features which are accurate in these tasks. The proposed hybrid classification framework combines a Support Vector Machine (SVM) and a Convolutional Neural Network (CNN). The SVM model uses feature extraction based on the orthogonal transformation (DCT), and the CNN model uses machine learning to represent data from Log-Mel spectral diagrams. Noisy sound signals are generated using the TIMIT and NOISEX-92 datasets by combining the clean signal with various types of environmental sounds, resulting in 17 categories, including the clean signal category. A soft voting strategy is then used to enhance overall decision-making. It calculates the weighted average probability for each class across all models and selects the class that has the highest average probability as the final output with more robust and reliable classification. The experimental results prove that the proposed model achieves high performance compared to current ways and individual models, reaching an accuracy of 99.28% with improvements in F1 score, accuracy and recall. The results obtained prove the efficiency of integrating ML and DL methods for classifying environmental sounds.