Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对零样本异常检测中正负样本语义重叠问题,提出Proximity-CLIP方法,通过视觉校准语义间距和异常查询模块提升检测性能。
📝 Abstract
Vision-language models offer a promising approach for zero-shot anomaly detection (ZSAD). However, due to object-centric bias, normal and anomalous text prototypes exhibit a high semantic overlap. While enforcing strict orthogonality between them improves discriminability, mapping highly contiguous visual inputs onto drastically orthogonal prototypes introduces a geometric dilemma, disrupting the pre-trained structural continuity. To address this problem, we propose Proximity-CLIP, a framework that visually calibrates the semantic margin to guide visual adaptation. First, we introduce a visually-calibrated semantic proximity learning mechanism that uses a bounded dynamic regularization to learn an appropriate semantic margin, ensuring discriminative separation while preserving structural alignment. Second, we design an Anomaly Query Module (AQM) driven by these text priors. Using the calibrated anomalous prototype as a semantic query, the AQM actively retrieves localized defect cues from contextual visual patches, mitigating the dilution of subtle anomalies during global pooling. Extensive experiments demonstrate that Proximity-CLIP outperforms current state-of-the-art methods across multiple ZSAD benchmarks with minimal architectural modifications.
Problem

Research questions and friction points this paper is trying to address.

zero-shot anomaly detection
semantic overlap
orthogonality
geometric dilemma
pre-trained structural continuity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proximity-CLIP
Semantic Proximity Learning
Anomaly Query Module (AQM)
Zero-Shot Anomaly Detection (ZSAD)
🔎 Similar Papers
2023-10-29International Conference on Learning RepresentationsCitations: 114
💼 Related Jobs
No related jobs found.
M
Manwen Yang
School of Software Engineering, Xi’an Jiaotong University; State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
L
Leqian Ding
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University
Yu Guo
Yu Guo
Xi’an Jiaotong University
6D pose estimationtime series predictiongraph learning
F
Fei Wang
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University