Institution profile

MVTec Software GmbH

Industry researcheurope · de
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors

Oct 02, 2025

Industrial safety inspection faces challenges in few-shot, class-agnostic anomaly detection, where distinguishing normal from anomalous features is inherently difficult due to subtle and diverse anomalies. Method: We propose a lightweight embedding-difference-driven approach that exploits the strong correlation between anomaly severity and local discrepancies in the embedding space of pretrained vision encoders (e.g., DINOv3). Instead of introducing complex architectures or requiring auxiliary supervision, we design a learnable nonlinear projection operator to explicitly unlock the implicit anomaly discriminability embedded in these representations. Our method models the natural image distribution on the embedding manifold using only a few normal samples and localizes out-of-distribution anomalies via difference heatmaps. Results: The approach achieves state-of-the-art performance across multiple industrial anomaly detection benchmarks, reduces model parameters by an order of magnitude, exhibits strong cross-category generalization, and demonstrates consistent effectiveness across diverse foundation encoders.

0 citationsRead paper

From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge

Sep 22, 2025

This study addresses two core challenges in real-world visual anomaly detection: robustness to distributional shifts and few-shot generalization. Leveraging the VAND 3.0 challenge, we propose a novel evaluation framework that jointly emphasizes out-of-distribution robustness and few-shot adaptability. We systematically investigate large-scale pre-trained vision and vision-language models (e.g., ViT, CLIP), introducing a backbone-driven pipeline for feature fusion and few-shot adaptation—incorporating optimized fine-tuning strategies and cross-modal feature alignment. Experiments demonstrate substantial improvements over baselines on both tasks, validating the critical contribution of advanced backbone architectures. Furthermore, our analysis uncovers a fundamental trade-off between computational efficiency and real-time deployability, highlighting bottlenecks in current approaches. The work thus provides new insights, practical design principles, and a rigorous benchmark for lightweight, robust, and production-ready anomaly detection systems.

0 citationsRead paper

MVTOP: Multi-View Transformer-based Object Pose-Estimation

Aug 05, 2025

In multi-view rigid object pose estimation, single-view methods suffer from pose ambiguity caused by occlusion and symmetry. To address this, we propose an end-to-end multi-view fusion framework grounded in line-of-sight (LOS) geometric modeling. Unlike conventional post-hoc fusion or depth-dependent approaches, our method integrates multi-view images early in feature extraction and incorporates camera intrinsics and relative pose priors to guide spatial reasoning via a novel LOS-aware attention mechanism. Crucially, it achieves global multi-view pose estimation without requiring depth supervision—a first in the literature. Extensive experiments on our synthetic dataset and the YCB-Video benchmark demonstrate significant improvements over both single-view baselines and state-of-the-art multi-view methods, validating robustness under ambiguous conditions and strong generalization capability.

0 citationsRead paper

Challenges of Requirements Communication and Digital Assets Verification in Infrastructure Projects

Apr 29, 2025

This study addresses critical challenges in requirements communication and digital asset verification between clients and suppliers in infrastructure projects. Drawing on two domain-specific case studies from road and railway engineering, we conducted semi-structured interviews with ten subject-matter experts and applied thematic coding analysis. For the first time, we systematically identified 13 recurrent challenges, analyzing their root causes, operational impacts, and striking parallels with well-documented software engineering problems. Our contribution lies in advancing cross-domain solution transfer: we synthesize reusable collaboration mechanisms and verification pathways grounded in empirical evidence. These findings provide both theoretical grounding and practical foundations for developing standardized, trustworthy frameworks for requirements coordination and digital asset validation in infrastructure delivery. (128 words)

0 citationsRead paper

BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation

Apr 03, 2025

This work addresses two key challenges in real-world 6D object pose estimation: model-free pose estimation (i.e., without requiring 3D CAD models) and open-set 6D detection under unknown object identities. We propose FreeZeV2.1, a unified framework integrating reference-video-guided model-free learning, cross-modal feature alignment, lightweight cooperative (Co-op) inference, and diffusion-enhanced pose regression. Additionally, we design MUSE, a vision-language-driven detector. We formally define the novel tasks of model-free pose estimation and open-set 6D detection, and introduce BOP-H3—the first high-fidelity multimodal benchmark featuring AR/VR-captured videos and corresponding 3D models. Experiments show that FreeZeV2.1 achieves a 22% absolute accuracy gain on BOP-Classic-Core; Co-op inference accelerates inference by 25× while improving accuracy by 13%; and MUSE yields a 21% relative improvement in 2D detection for unseen objects. Code and an online evaluation platform are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors

Oct 02, 2025

Industrial safety inspection faces challenges in few-shot, class-agnostic anomaly detection, where distinguishing normal from anomalous features is inherently difficult due to subtle and diverse anomalies. Method: We propose a lightweight embedding-difference-driven approach that exploits the strong correlation between anomaly severity and local discrepancies in the embedding space of pretrained vision encoders (e.g., DINOv3). Instead of introducing complex architectures or requiring auxiliary supervision, we design a learnable nonlinear projection operator to explicitly unlock the implicit anomaly discriminability embedded in these representations. Our method models the natural image distribution on the embedding manifold using only a few normal samples and localizes out-of-distribution anomalies via difference heatmaps. Results: The approach achieves state-of-the-art performance across multiple industrial anomaly detection benchmarks, reduces model parameters by an order of magnitude, exhibits strong cross-category generalization, and demonstrates consistent effectiveness across diverse foundation encoders.

0 citationsRead paper

From Benchmarks to Reality: Advancing Visual Anomaly Detection by the VAND 3.0 Challenge

Sep 22, 2025

This study addresses two core challenges in real-world visual anomaly detection: robustness to distributional shifts and few-shot generalization. Leveraging the VAND 3.0 challenge, we propose a novel evaluation framework that jointly emphasizes out-of-distribution robustness and few-shot adaptability. We systematically investigate large-scale pre-trained vision and vision-language models (e.g., ViT, CLIP), introducing a backbone-driven pipeline for feature fusion and few-shot adaptation—incorporating optimized fine-tuning strategies and cross-modal feature alignment. Experiments demonstrate substantial improvements over baselines on both tasks, validating the critical contribution of advanced backbone architectures. Furthermore, our analysis uncovers a fundamental trade-off between computational efficiency and real-time deployability, highlighting bottlenecks in current approaches. The work thus provides new insights, practical design principles, and a rigorous benchmark for lightweight, robust, and production-ready anomaly detection systems.

0 citationsRead paper

MVTOP: Multi-View Transformer-based Object Pose-Estimation

Aug 05, 2025

In multi-view rigid object pose estimation, single-view methods suffer from pose ambiguity caused by occlusion and symmetry. To address this, we propose an end-to-end multi-view fusion framework grounded in line-of-sight (LOS) geometric modeling. Unlike conventional post-hoc fusion or depth-dependent approaches, our method integrates multi-view images early in feature extraction and incorporates camera intrinsics and relative pose priors to guide spatial reasoning via a novel LOS-aware attention mechanism. Crucially, it achieves global multi-view pose estimation without requiring depth supervision—a first in the literature. Extensive experiments on our synthetic dataset and the YCB-Video benchmark demonstrate significant improvements over both single-view baselines and state-of-the-art multi-view methods, validating robustness under ambiguous conditions and strong generalization capability.

0 citationsRead paper

Challenges of Requirements Communication and Digital Assets Verification in Infrastructure Projects

Apr 29, 2025

This study addresses critical challenges in requirements communication and digital asset verification between clients and suppliers in infrastructure projects. Drawing on two domain-specific case studies from road and railway engineering, we conducted semi-structured interviews with ten subject-matter experts and applied thematic coding analysis. For the first time, we systematically identified 13 recurrent challenges, analyzing their root causes, operational impacts, and striking parallels with well-documented software engineering problems. Our contribution lies in advancing cross-domain solution transfer: we synthesize reusable collaboration mechanisms and verification pathways grounded in empirical evidence. These findings provide both theoretical grounding and practical foundations for developing standardized, trustworthy frameworks for requirements coordination and digital asset validation in infrastructure delivery. (128 words)

0 citationsRead paper

BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation

Apr 03, 2025

This work addresses two key challenges in real-world 6D object pose estimation: model-free pose estimation (i.e., without requiring 3D CAD models) and open-set 6D detection under unknown object identities. We propose FreeZeV2.1, a unified framework integrating reference-video-guided model-free learning, cross-modal feature alignment, lightweight cooperative (Co-op) inference, and diffusion-enhanced pose regression. Additionally, we design MUSE, a vision-language-driven detector. We formally define the novel tasks of model-free pose estimation and open-set 6D detection, and introduce BOP-H3—the first high-fidelity multimodal benchmark featuring AR/VR-captured videos and corresponding 3D models. Experiments show that FreeZeV2.1 achieves a 22% absolute accuracy gain on BOP-Classic-Core; Co-op inference accelerates inference by 25× while improving accuracy by 13%; and MUSE yields a 21% relative improvement in 2D detection for unseen objects. Code and an online evaluation platform are publicly released.

0 citationsRead paper