科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on pattern analysis and machine intelligence2026-09-09

Enhancing Local Cognition of CLIP for Training-Free Open Vocabulary Semantic Segmentation.

Tong Shao, Zhuotao Tian, Yunyang Mo, Guoying Sun, Li Jiang, Hengshuang Zhao, Jingyong Su

原始摘要(英文原文)· Original abstract
CLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot generalization. Despite its success, its application to OVSS is limited due to its initial image-level alignment training, which affects its performance in tasks requiring detailed local context. Our study delves into the impact of CLIP's [CLS] token on patch feature correlations, revealing a dominance of the "global patch" that hinders local feature discrimination and demonstrating the reasons, trends and impacts of its occurrence. To overcome this, we propose CLIPtrase, a novel training-free semantic segmentation strategy, which enhances local cognition through recalibrated self-correlation among patches and quadrant mask ensemble strategy. This approach demonstrates significant improvements in segmentation accuracy and the ability to maintain semantic coherence between objects. The experiments indicate a 22.8% average improvement over CLIP in 9 segmentation benchmarks, surpassing current state-of-the-art training-free methods. Moreover, the proposed method can be seamlessly adapted to other tasks such as unsupervised semantic segmentation and multi-label classification with satisfactory performance, which demonstrate the generalization of our approach. The code is available at https://github.com/leaves162/CLIPtrase.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Enhancing Local Cognition of CLIP for Training-Free Open Vocabulary Semantic Segmentation. — 科研速览 Science Skim