Guangzeng Han, James G Murphy, Benjamin O Ladd, Xiaolei Huang, Brian Borsari
Multimodal self-consistency outperforms single-pass baseline prompting approaches. These findings suggest that incorporating both what clients say and how they say it can support more reliable MI coding.
INTRODUCTION: Understanding client behaviors and predicting their outcomes require coding Motivational interviewing (MI) sessions with intensive labor costs and time consumptions of MI professionals. The advances in audio language models (ALMs) open promising opportunities in automating coding process, capturing critical multimodal signals of behavioral patterns, and interactive collaborations with expert annotators.
MATERIALS AND METHODS: We experimented with 5 recorded sessions from de-identified MI audio tapes. We deployed audio-language models with four complementary analytic prompts to augment utterance-level reasoning: analytic (verbal cues), prosody-aware (acoustic cues), evidence-scoring (quantitative hypothesis test), and comparative (contrastive reasoning). Three stochastic samples were drawn per prompt, generating 12 independent reasoning trajectories per utterance, and majority votes of all trajectories decided predictions.
RESULTS: Performance was evaluated using accuracy, precision, recall, and macro-F1 scores. The proposed MM-SC approach achieved 52.56% accuracy, 54.03% precision, 47.45% recall, and a macro-F1 score of 46.40%, exceeding the performance of baseline methods. Systematic ablation removing individual modules consistently degraded performance on primary metrics.
CONCLUSIONS: Multimodal self-consistency outperforms single-pass baseline prompting approaches. These findings suggest that incorporating both what clients say and how they say it can support more reliable MI coding.