Yunrui Cai
This project investigates the automatic understanding of multi-speaker conversations, with the aim of identifying who spoke, what was said, when it was said, and what topics were discussed. The research will develop a unified framework combining speech recognition, speaker diarization, timestamp alignment, and topic segmentation, with support from text and audio prompts. The project will use existing and generated multi-speaker audio recordings, speech transcripts, speaker labels, timestamps, topic-boundary annotations, acoustic features, and model outputs. These data will be used to train and evaluate models under conditions such as overlapping speech, speaker changes, and long conversations. The expected outcome is a more robust and controllable system for applications including meeting transcription, dialogue summarization, information retrieval, and conversational analysis.