Vijay Ravi, Camille Noufi
Conversational primary care audio contains an acoustic signal that discriminates clinically meaningful respiratory contrasts. Absolute performance is moderate, but the conditions are stricter than in prior work: conversational speech and differential-diagnosis contrasts among patients with illness. This pilot study establishes a baseline for voice-based clinical AI, advancing from sick-vs.-healthy detection toward differential-diagnosis panels and demonstrating a proof of concept for hierarchical composition.
BACKGROUND: Respiratory complaints account for a substantial share of adult ambulatory visits, and accurate triage has direct consequences for antibiotic stewardship and pathogen-specific therapy. Prior work has investigated voice as a triage signal, but that literature is dominated by single-condition detection from scripted speech in crowdsourced or controlled clinical settings and has not been evaluated at the primary care scale using conversational ambient audio.
METHODS: A dataset of 514,377 ambient-recorded primary care visits from 379,225 adult patients at a US clinic network was used, with per-visit clinically assigned ICD-10 diagnosis codes and de-identified demographic and geographic metadata. Patient audio was extracted from each doctor-patient conversation, and spectral, voice quality, and prosodic features were computed. Eleven binary classification tasks were defined, aligned with a respiratory triage cascade (e.g., acute respiratory vs. acute non-respiratory illness, and lower vs. upper respiratory tract infection). An acoustic model was trained independently for each task using patient-stratified 5-fold cross-validation and evaluated on a held-out test set. Each model was also compared against six non-acoustic baselines using a single demographic, geographic, or temporal variable. The 11 trained classifiers were combined into a hierarchical cascade and illustrated as case studies.
RESULTS: Test-set AUC across the 11 tasks ranged from 0.602 (95% CI: 0.588-0.614) to 0.745 (95% CI: 0.742-0.748), with a mean expected calibration error of 0.018. After multiple-testing correction, six of the eleven binaries outperformed all six confounder baselines. Four binaries showed a median within-stratum AUC of 0.61-0.70 when the confounder was held fixed, indicating acoustic discrimination beyond what the confounder alone explains. Five binaries failed at least one axis; the only one outperformed by a confounder baseline was the pneumonia vs. non-pneumonia lower respiratory tract infection binary, which failed against the patient-city confounder baseline, plausibly reflecting a clinic-level difference in ICD-10 coding.
CONCLUSION: Conversational primary care audio contains an acoustic signal that discriminates clinically meaningful respiratory contrasts. Absolute performance is moderate, but the conditions are stricter than in prior work: conversational speech and differential-diagnosis contrasts among patients with illness. This pilot study establishes a baseline for voice-based clinical AI, advancing from sick-vs.-healthy detection toward differential-diagnosis panels and demonstrating a proof of concept for hierarchical composition.