Tasnova Haque Mazumder, Akash Ahmed, Dilruba Akter, Amin Ocin, Raihan Ul Islam, Mohammad Rifat Ahmmad Rashid, Shamim H Ripon, Ahmed Wasif Reza
Environmental sound classification (ESC) has been used to implement various applications such as smart city, environmental monitoring systems, context-aware systems, and assistive technologies. The existing benchmark datasets, however, mainly focus on the acoustic environment of the west and limited acoustic scenes are available in Bangladesh. In this paper, we introduce a real-world environmental audio dataset named AcousticSceneBD, with a total of 5535 recordings across 7 classes of acoustic scenes: Bus, Metro, Metro Station, Park, Restaurant, Shopping Mall, and University. All recordings were conducted using the standard operating conditions of the smartphone's microphone and with all files standardized following the same procedure:10-second, mono, 16-bit WAV audio recorded at 44.1/48 kHz, with a total duration of 15.38 hours and a dataset size of 4.87 GB.To guarantee consistency and reproducibility, the dataset has been created using a structured approach for data collection and processing consisting of eight stages. Four baseline models, namely Random Forest, YAMNet, Wav2Vec2, and PANNs CNN14 were tested to validate the data, with respective test accuracy of 58.81%, 82.99%, 47.11%, and 78.15%, where Random forest is trained on extracted MFCC features from the raw audio The results indicate that AcousticSceneBD is well annotated, acoustically discriminative and appropriate for the development and evaluation of traditional, deep learning, transfer learning, and self-supervised environmental sound classification methods.