科研速览 · Science Skim继续刷下去 · Keep skimming →
◇ bioRxiv2026-09-03· neuroscience

Human-like meaning maps from single-prompt VLM ratings of local scene meaning

K. Jurewicz, Y.-E. Berreby, S. Krishna

原始摘要(英文原文)· Original abstract
Visual attention is shaped both by low-level perceptual features and by the semantic properties of scene regions. However, compared to the salience of low-level image features, the characterization of image properties that lead to scene understanding has been more difficult. The semantic properties of scene regions have been operationalized via meaning maps, proposed by Henderson and Hayes (2017), by having human observers rate the informativeness of scene patches. However, this rating procedure is labor-intensive and limits the method's versatility. An automated alternative, the DeepMeaning model, reduces rating costs but still requires human training data from a similar class of images, limiting generalizability. Here, we tested whether instruction-following vision-language models (VLMs) can generate human-like ratings in a zero-shot manner: using the original human-rating instructions as a single prompt without any task-specific training. We obtained meaningfulness ratings from both open-weight and proprietary VLMs for scene patches across three meaning-map datasets, and compared the resulting zero-shot AI meaning maps (AIMMs) to human meaning maps and maps obtained from the DeepMeaning model. Zero-shot AIMMs closely approximate human meaning maps, and maps from the task-specific DeepMeaning model, showing that VLMs can replicate human semantic judgments without any training. Zero-shot AIMMs demonstrate utility and validity, matching human meaning maps in their ability to predict human fixations and reproducing the same pattern of relative performance against other gaze-prediction models. We provide open-source code that allows researchers to use this zero-shot approach to quickly and flexibly explore meaning maps across spatial scales, prompts, and scene categories.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Human-like meaning maps from single-prompt VLM ratings of local scene meaning — 科研速览 Science Skim