DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection
This work proposes a fully automatic, training-free few-shot anomaly detection framework that addresses the limitation of existing methods, which predominantly rely on local image patch features while overlooking the global contextual information embedded in the [CLS] token of Vision Transformers. The study is the first to reveal and exploit the dual nature of the [CLS] token: its global semantic invariance and its attention map’s ability to indicate spatial anomalies. By integrating a semantic consistency-driven automatic augmentation strategy with an attention-guided dynamic feature reweighting mechanism, the method achieves precise anomaly localization and scoring without manual hyperparameter tuning. Under single-sample settings on MVTec-AD, VisA, and Real-IAD, it attains Image-AUC scores of 97.7%, 93.2%, and 84.5%, respectively, demonstrating plug-and-play state-of-the-art performance across categories, backbone architectures, and datasets.