科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Current biology : CB2026-08-31

Shared texture-like representations underlie deep neural network alignment with human visual processing.

Jessica Loke, Lynn K A Sörensen, Iris I A Groen, Natalie Cappaert, H Steven Scholte

原始摘要(英文原文)· Original abstract
Deep neural networks (DNNs) excel at predicting neural responses across the visual hierarchy,1,2,3,4,5 a success widely interpreted as evidence of shared object recognition computations.6,7 Yet improving DNN object recognition accuracy does not reliably increase neural predictivity,8,9 and even untrained networks predict brain responses above chance.9,10,11 This disconnect suggests that object recognition may not drive DNN-brain alignment. Texture-like statistics are represented in both DNNs and mid-level visual cortex (A.V. Jagadeesh and M. Livingstone, 2024, ICLR, presentation).12,13,1416 In natural images, these statistics are carried by objects and backgrounds, shaping representations and recognition in both systems.16,17,18,19 Does DNN-brain alignment reflect a shared sensitivity to object-related information or texture-like statistics? To dissociate these factors, we recorded electroencephalograms (EEGs) from 57 participants viewing natural scenes, texture-synthesized images preserving local statistics while disrupting global form, and object-only images with backgrounds removed. If alignment reflects texture-like statistics, then it should peak for texture-synthesized images. If it reflects object-related processing, then alignment should be strongest for natural and object-only conditions, which preserve object information. We compared EEG responses with DNN activations via weighted representational similarity analysis.20,21 Texture-synthesized images yielded the strongest DNN-EEG alignment, peaking in early responses (<200 ms) and explaining up to ∼85% of noise-ceiling-normalized explainable variance versus ∼44% for natural and ∼55% for isolated objects. Crucially, object categories were more decodable for natural and object-only images than texture-synthesized images, yet these object-rich conditions showed weaker alignment. This dissociation reveals that DNNs capture the texture-statistical component of early visual responses while failing to explain later, object-related variance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Shared texture-like representations underlie deep neural network alignment with human visual processing. — 科研速览 Science Skim