科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Transactions on Multimedia2026-01-01· Computer science

Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph Captioning

Shuqin Chen, Xian Zhong, Xingrui Yang, Li Yang, Bo Sheng, Alex C. Kot

原始摘要(英文原文)· Original abstract
Video paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments onActivityNet CaptionsandYouCook2demonstrate FLS-Net's superiority over state-of-the-art approaches. The source code is available athttps://github.com/yangxingrui/FLS.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph Captioning — 科研速览 Science Skim