Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
Industrial safety inspection faces challenges in few-shot, class-agnostic anomaly detection, where distinguishing normal from anomalous features is inherently difficult due to subtle and diverse anomalies. Method: We propose a lightweight embedding-difference-driven approach that exploits the strong correlation between anomaly severity and local discrepancies in the embedding space of pretrained vision encoders (e.g., DINOv3). Instead of introducing complex architectures or requiring auxiliary supervision, we design a learnable nonlinear projection operator to explicitly unlock the implicit anomaly discriminability embedded in these representations. Our method models the natural image distribution on the embedding manifold using only a few normal samples and localizes out-of-distribution anomalies via difference heatmaps. Results: The approach achieves state-of-the-art performance across multiple industrial anomaly detection benchmarks, reduces model parameters by an order of magnitude, exhibits strong cross-category generalization, and demonstrates consistent effectiveness across diverse foundation encoders.