科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Pattern Recognition2026-04-12· Closed captioning

GroundCap: A visually grounded image captioning dataset with object and action identification

Daniel A. P. Oliveira, Lourenço Teodoro, David Martins de Matos

原始摘要(英文原文)· Original abstract
Current image captioning systems lack the ability to link descriptive text to specific visual elements, making their outputs difficult to verify. While recent approaches offer some grounding capabilities, they cannot track object identities across multiple references or ground both actions and objects simultaneously. We propose a novel ID-based grounding system that enables consistent object reference tracking and action–object linking. We present GroundCap, a dataset containing 52,016 images from 77 movies, with 334 human-annotated and 52,016 automatically generated captions. Each caption is grounded on detected objects (132 classes) and actions (51 classes) using a tag system that maintains object identity while linking actions to the corresponding objects. Our approach features persistent object IDs for reference tracking, explicit action–object linking, and the segmentation of background elements through K-means clustering. We propose gMETEOR, a composite metric that jointly measures grounding accuracy and language quality, and establish baseline performance by fine-tuning Pixtral-12B and Qwen2.5-VL 7B on GroundCap. Human evaluation demonstrates our approach’s effectiveness in producing verifiable descriptions with coherent object references.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

GroundCap: A visually grounded image captioning dataset with object and action identification — 科研速览 Science Skim