Yun Peng, Xiao Lin, Nachuan Ma, Chengju Liu, Qijun Chen
Logical anomaly detection aims to identify inconsistencies in object relationships and scene semantics, which is a challenging task requiring high-level reasoning capabilities that go beyond low-level structural cues. Existing methods rely heavily on large amounts of labeled data, struggle in zero-shot scenarios and lack deep semantic understanding in complex scenes with multiple objects. To address this, we propose VLLM-LAD, a visual large language model framework specifically designed for zero-shot logical anomaly detection. It incorporates a Local Detail Visual Perceiver (LDVP) module for fine-grained anomaly localization and a Global Semantic Feature Fusion (GSFF) module for enhanced multimodal reasoning. We also introduce a global-local-regional contrastive learning strategy to align visual and textual representations, thereby achieving robust zero-shot performance. Extensive experiments on MVTec LOCO demonstrate that VLLM-LAD outperforms existing methods in accuracy and interpretability. To further advance research in this area, we introduce the RLAD dataset. This large-scale, diverse dataset is designed for logical anomaly detection, encompassing complex scenarios that demand reasoning about object relationships and global semantics. Our method not only excels on the RLAD dataset but also provides detailed textual explanations and segmentation masks, enhancing its transparency and applicability to real-world tasks. This work establishes a new benchmark for zero-shot logical anomaly detection, offering a robust framework and a valuable dataset to propel future research.