Himashi Rathnayake, Jesin James, G. Leoni, Ake Nicholas, Catherine Watson, Peter Keegan
Speech emotion recognition (SER) is an emerging field in human–computer interaction. Although numerous studies have focused on SER for well-resourced languages, the literature reveals a significant gap in research on low-resource and Indigenous (LRI) languages. This paper presents a comprehensive review of the existing literature on SER in the context of LRI languages, analysing critical factors to consider at each stage of designing an SER system. The review indicates that most studies on SER for LRI languages adopt emotion categories established for well-resourced languages, often assuming the universality of emotions. However, the literature suggests that this approach may be limited due to emotional disparities influenced by cultural variations. Additionally, the review underscores that current SER systems typically lack community-oriented methodologies in the development of technology for LRI languages. The importance of feature selection is highlighted, with evidence suggesting that a combination of traditional machine learning methods and carefully selected acoustic features may offer viable options for SER in these languages. Furthermore, the review identifies a need for further exploration of semi-supervised and unsupervised approaches to enhance SER capabilities in LRI contexts. Overall, current SER systems for LRI languages lag behind state-of-the-art standards due to the lack of resources, indicating that there is still much work to be done in this area.