AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
Zero-shot anomaly detection (ZSAD) aims to identify anomalies across domains without access to target-domain training samples; however, its generalizability is severely hindered by substantial discrepancies in foreground objects, anomaly appearances, and background distributions. To address this, we propose a CLIP-based universal ZSAD framework. Our method introduces the first object-agnostic text prompt learning mechanism, decoupling foreground semantics from normal/abnormal discrimination modeling. We further design learnable, domain-agnostic “normal” and “abnormal” text prompts and leverage vision-language feature alignment to enable zero-shot anomaly scoring and pixel-level localization. Critically, our approach requires no target-domain annotations or fine-tuning. Extensive experiments across 17 industrial defect and medical imaging datasets demonstrate significant improvements over existing ZSAD methods, achieving both strong cross-domain generalization and high localization accuracy.