科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE journal of biomedical and health informatics2026-09-03

V2D: Detection of type 2 diabetes using only smartphone voice recordings via spectro-temporal transformer embeddings.

Boujemaa Guermazi, Alexandre Logerais, Erica Y Huynh, Jessica Oreskovic, Adam Callanan, Yan Fossat

原始摘要(英文原文)· Original abstract
Voice-based screening offers a noninvasive, scalable avenue for early detection of type 2 diabetes using everyday smartphone recordings and acoustic features alone. We present V2D (Voice2Diabetes), a novel application of spectrogram-transformer embeddings derived exclusively from short speech segments for patient-level diabetes classification, without requiring any clinical measures or demographic variables. Adults (n=461; 157 female, 304 male) each completed multiple smartphone recordings while reading randomly selected sentences on their own smartphones. Mel-spectrograms were encoded with a pretrained Audio Spectrogram Transformer (AST) to train sex-stratified patient-level classifiers under fivefold nested cross-validation with a held-out calibration set; predictions were aggregated across recordings per participant. Using acoustic features alone, the models achieved patient-level balanced accuracy of 0.724 ± 0.023 (males) and 0.713 ± 0.021 (females), with area under the receiver operating characteristic curve (AUC) of 0.779 ± 0.019 and 0.788 ± 0.018, respectively, averaged over five independent random seeds. The pipeline incorporated probability calibration, sensitivity-first threshold optimization, and optional interpretable acoustic anchors (e.g., fundamental frequency, harmonic-to-noise ratio) to support clinical interpretation. These results provide a rigorous technical validation of AST-derived embeddings for acoustic-only, sex-stratified T2D classification in smartphone recordings and motivate prospective external validation in broader populations.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

V2D: Detection of type 2 diabetes using only smartphone voice recordings via spectro-temporal transformer embeddings. — 科研速览 Science Skim