科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ arXiv2026-09-13· eess.AS

Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection

Willem Berner, Julio Cesar Cavalcanti, Kalle Åström, Gabriel Skantze

原始摘要(英文原文)· Original abstract
Turn-taking is a fundamental component of spoken interaction, and while humans naturally rely on both verbal and non-verbal signals, dialogue systems usually depend on audio cues alone. This paper investigates whether visual features from face-to-face conversations can enhance turn-taking prediction beyond what is achievable from audio-only. We extend the Voice Activity Projection (VAP) model, a self-supervised transformer-based model for predicting future voice activity, by incorporating visual features extracted from the large-scale Meta Seamless Interaction dataset of dyadic face-to-face conversations. The visual features include gaze direction, head movement, body and hand pose, and facial action units (FAU). For incorporating the visual features, we explore concatenation, cross-attention fusion, delta features, and trainable gating mechanisms. Results show that visual information improves performance over the audio-only baseline, with FAU being significantly more informative than other feature groups. Body and gaze features nevertheless contribute complementary information, as the model combining all features performs best. Furthermore, results indicate that performance on specific tasks varies depending on whether training and test data come from improvised (acted) or naturalistic (non-acted) conversations.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Exploring Multimodal Turn-Taking Cues in Face-to-Face Conversation using Voice Activity Projection — 科研速览 Science Skim