科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR aging2026-09-21

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image.

Felix Agbavor, Hualou Liang

一句话结论 · In one sentence

These results demonstrate that attention-based multimodal fusion can enhance dementia prediction from picture-description responses and provide a strong foundation for developing multimodal cognitive screening pipelines.

原始摘要(英文原文)· Original abstract
BACKGROUND: Alzheimer disease (AD) is a leading cause of dementia, and there is growing interest in scalable approaches for early screening using speech-based tasks. While prior work has demonstrated promising results using either transcript-based language features or acoustic cues, most approaches remain unimodal or rely on simple fusion strategies that do not explicitly consider interactions across modalities. OBJECTIVE: In this study, we propose an attention-based trimodal fusion framework that integrates text, audio, and image representations of the Cookie Theft picture, which serves as the shared visual stimulus in the picture-description task. METHODS: Our method uses a new bidirectional cross-attention mechanism to achieve a unified multimodal embedding for downstream tasks. We evaluate the approach on 2 tasks: AD detection by classifying whether the participant has AD or not, and AD severity assessment by predicting Mini-Mental Status Examination cognitive scores. RESULTS: On the AD detection task, trimodal fusion achieves the best overall performance (F1-score=0.8667, area under the receiver operating characteristic curve=0.9032), outperforming unimodal baselines, bimodal fusion, and conventional early or late fusion methods. For AD severity assessment, the proposed multimodal representation reduces prediction error of root mean squared error to about 4.20, improving over both unimodal and bimodal fusion settings. We further perform the ablation analysis to show that bidirectional cross-attention consistently outperforms conventional unidirectional cross-attention. CONCLUSIONS: These results demonstrate that attention-based multimodal fusion can enhance dementia prediction from picture-description responses and provide a strong foundation for developing multimodal cognitive screening pipelines.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image. — 科研速览 Science Skim