Pengxu Jiang, Wenke Zhang, Aiqin Li, Peng Li
These results suggest that sustained vowel phonations contain potentially informative acoustic patterns associated with T2DM status. However, further validation using independent cohorts is required before clinical application.
PROBLEM: While voice-based analysis has emerged as a potential non-invasive approach for exploring acoustic alterations associated with Type 2 Diabetes Mellitus (T2DM), current approaches often fail to fully capture disease-relevant vocal patterns due to simplified recording protocols and insufficient feature modeling.
AIM: This study aims to investigate the feasibility of voice-based classification of T2DM using a multi-feature fusion framework, leveraging sustained vowel phonations and integrating multiple acoustic feature types to enhance diagnostic performance.
METHODS: Voice recordings were collected from 378 participants, including T2DM patients and healthy controls. Each participant pronounced six Mandarin vowels, from which acoustic features-including Mel-frequency cepstral coefficients (MFCCs), glottal parameters, and eGeMAPS descriptors-were extracted. Neural networks, including convolutional and deep architectures, were applied to capture pathology-relevant segments. Vowel-level fusion modules aggregated features across vowels, and an attention-based hierarchical fusion mechanism combined the three feature types, adaptively emphasizing features most relevant to T2DM.
RESULTS: The proposed dual-branch hierarchical fusion model achieved 80.48% detection accuracy.
CONCLUSION: These results suggest that sustained vowel phonations contain potentially informative acoustic patterns associated with T2DM status. However, further validation using independent cohorts is required before clinical application.