Stephen Ojo, Nihal Abuzinadah, Ghada Atteia, Khaled Alnowaiser, Abeer Aljohani, Muhammad Umer, Refka Ghodhbani, Thomas I Nathaniel
Accurate gender and age classification from voice data is important for personalized, secure, and ethical human-computer interaction (HCI). With the growing use of large pretrained speech models such as wav2vec 2.0, there is a need to leverage these representations responsibly for multilingual and privacy-aware demographic inference. This study proposes VoiceNet-RAI, a deep learning (DL) framework that integrates Mel-Frequency Cepstral Coefficients (MFCCs) with transformer-based multilingual wav2vec 2.0 embeddings within an EfficientNet-Lite architecture to support accurate, fair, and transparent classification across diverse linguistic contexts. The framework combines contextual and spectral information to improve robustness under variable recording conditions. Experiments on multilingual speech data covering 28 languages demonstrate strong classification performance, with the English subset achieving 97.87% accuracy, 98.61% precision, 98.70% recall, and a 98.65% F1-score. External validation using the Mozilla Common Voice dataset yielded 98.27% accuracy for gender classification, demonstrating generalizability to an independent dataset. Ablation analysis further showed that combining MFCC and wav2vec 2.0 features improved performance compared with raw and MFCC-only representations. These findings demonstrate the potential of hybrid spectral-contextual speech representations for robust and responsible gender and age classification.