🤖 AI Summary
Existing benchmarks such as MVTec AD and VisA exhibit saturation in the AU-PRO metric, resulting in low model discriminability and hindering progress in industrial anomaly detection. To address this, we introduce MVTec AD 2—the first benchmark specifically designed for high-difficulty industrial scenarios. It comprises 8 challenging task categories (e.g., transparent/occluded objects, dark-field imaging, high-variance normal samples, sub-pixel defects) and over 8,000 high-resolution images. Crucially, it is the first to systematically incorporate realistic distribution shifts—such as illumination variations—to rigorously evaluate model robustness. All anomalies are annotated at pixel-level precision, and a standardized evaluation protocol—including a publicly accessible evaluation server—is provided. State-of-the-art methods achieve an average AU-PRO below 60%, substantially widening performance gaps and effectively breaking the AU-PRO saturation bottleneck. MVTec AD 2 thus establishes a new, reproducible, and highly discriminative benchmark for industrial anomaly detection research.
📝 Abstract
In recent years, performance on existing anomaly detection benchmarks like MVTec AD and VisA has started to saturate in terms of segmentation AU-PRO, with state-of-the-art models often competing in the range of less than one percentage point. This lack of discriminatory power prevents a meaningful comparison of models and thus hinders progress of the field, especially when considering the inherent stochastic nature of machine learning results. We present MVTec AD 2, a collection of eight anomaly detection scenarios with more than 8000 high-resolution images. It comprises challenging and highly relevant industrial inspection use cases that have not been considered in previous datasets, including transparent and overlapping objects, dark-field and back light illumination, objects with high variance in the normal data, and extremely small defects. We provide comprehensive evaluations of state-of-the-art methods and show that their performance remains below 60% average AU-PRO. Additionally, our dataset provides test scenarios with lighting condition changes to assess the robustness of methods under real-world distribution shifts. We host a publicly accessible evaluation server that holds the pixel-precise ground truth of the test set (https://benchmark.mvtec.com/). All image data is available at https://www.mvtec.com/company/research/datasets/mvtec-ad-2.