科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ California Digital Library2026-07-31· Computer science

Unifying Multi-Speaker and Topic Segmentation with Multimodal Prompts

Yunrui Cai

原始摘要(英文原文)· Original abstract
This project investigates the automatic understanding of multi-speaker conversations, with the aim of identifying who spoke, what was said, when it was said, and what topics were discussed. The research will develop a unified framework combining speech recognition, speaker diarization, timestamp alignment, and topic segmentation, with support from text and audio prompts. The project will use existing and generated multi-speaker audio recordings, speech transcripts, speaker labels, timestamps, topic-boundary annotations, acoustic features, and model outputs. These data will be used to train and evaluate models under conditions such as overlapping speech, speaker changes, and long conversations. The expected outcome is a more robust and controllable system for applications including meeting transcription, dialogue summarization, information retrieval, and conversational analysis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Unifying Multi-Speaker and Topic Segmentation with Multimodal Prompts — 科研速览 Science Skim