科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in digital health2026-01-01

Multimodal text-audio sentiment in clinical aphasia speech using NLP.

Shamiha Binta Manir, Anai N Kothari, William L Gross, Priya Deshpande

一句话结论 · In one sentence

Under the recording-disjoint split, both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1, while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings, while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus, aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions, but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances, indicating that affective-tone interpretation in fragmented clinical speech is ambiguous.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Aphasia affects expressive and receptive communication and may influence the affective tone expressed during clinical speech tasks. This study presents an exploratory weakly supervised NLP analysis of positive/negative affective-tone proxies in AphasiaBank transcripts with paired audio. METHODS: We extracted sentence embeddings from DistilBERT ( e ∈ R 768 ) and recording-level acoustic summaries ( MFCC 13 , ZCR, RMS, spectral centroid, and spectral bandwidth; a ∈ R 17 ). Text and acoustic features were concatenated ( x = [ e ; a ] ∈ R 785 ) and classified using Random Forest models. Sentiment labels were generated using an SST-2-derived weak-supervision pipeline and should be interpreted as pseudo-labels rather than clinical ground truth. To evaluate modality contribution and potential leakage, we compared text-only, audio-only, and fused text-audio models under utterance-level and recording-disjoint splits. A small five-rater evaluation was used to examine human judgment alignment. RESULTS: Under the recording-disjoint split, both the text-only and fused text-audio models achieved 97.9% accuracy and 0.791 macro-F1, while the audio-only model achieved 55.4% accuracy and 0.388 macro-F1. These results indicate that classification performance was primarily driven by textual embeddings, while the recording-level acoustic summaries did not improve performance over text-only features. Across the pseudo-labeled corpus, aphasic utterances were more often labeled negative than control utterances. Age-stratified summaries showed subgroup variation in pseudo-label distributions, but these patterns were treated descriptively because labels were model-derived. Human-rater agreement was low for aphasic utterances, indicating that affective-tone interpretation in fragmented clinical speech is ambiguous. DISCUSSION: These findings should be interpreted as exploratory evidence about weakly supervised affective-tone proxies, not as validated clinical sentiment recognition. The results highlight both the promise of clinical NLP for aphasia discourse analysis and the need for independent human-labeled validation, utterance-aligned acoustic features, and careful control of domain, task, age, and topic bias.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal text-audio sentiment in clinical aphasia speech using NLP. — 科研速览 Science Skim