Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of semantic misalignment and architectural dependency caused by auxiliary modality fusion in multimodal zero-shot anomaly detection. We propose a plug-and-play auxiliary condition enhancement framework that establishes a new paradigm shifting from semantic fusion to conditional modulation. Specifically, this method leverages meta-learning to generate adaptive low-rank residuals and incorporates an uncertainty-aware spatial modulation mechanism to refine RGB features. This approach enables selective enhancement while preserving the original image-text matching pathway. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets demonstrate that the proposed framework significantly improves the performance of various detectors, achieving state-of-the-art results.
📝 Abstract
Recent foundation model-based methods have endowed RGB images with strong zero-shot anomaly detection (ZSAD) through vision-language pretraining. However, RGB observations alone remain limited in perceiving anomalies dominated by geometric deformation, depth variation, or subtle surface changes. Auxiliary modalities can provide complementary structural information, but existing multimodal methods typically fuse them directly into a shared semantic space, which may disturb the text-aligned anomaly semantics established by RGB foundation models and often requires modality-specific architectures. To address this issue, we propose a plug-and-play auxiliary-conditioned enhancement framework for zero-shot anomaly detection. Instead of reconstructing a joint multimodal anomaly semantic space, our framework preserves the original RGB image-text anomaly matching pathway and uses auxiliary observations as conditional signals for RGB feature refinement, allowing auxiliary modalities to seamlessly enhance existing RGB-based zero-shot anomaly detectors. Specifically, a lightweight meta-learning module takes global RGB and auxiliary representations as input and generates sample-adaptive low-rank residual updates to determine how RGB features should be refined. We further construct uncertainty-aware spatial modulation from the initial RGB anomaly response and auxiliary reliability, which determines where local residual updates are strengthened or suppressed. This global-to-local conditional modulation enables selective multimodal enhancement while preserving the original RGB anomaly semantics. Extensive experiments on MVTec 3D-AD and Eyecandies demonstrate that our framework consistently improves multiple popular RGB-based zero-shot anomaly detectors, achieving state-of-the-art performance for multimodal zero-shot anomaly detection.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot Anomaly Detection
Multimodal Learning
Auxiliary Modalities
Semantic Fusion
Foundation Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Modulation
Zero-shot Anomaly Detection
Meta-learning
Uncertainty-aware Spatial Modulation
Plug-and-play Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Peng Wu
School of Computer Science, Northwestern Polytechnical University, China
X
Xin Ge
School of Computer Science, Northwestern Polytechnical University, China
Yujia Sun
Yujia Sun
Xidian University
Deep LearningHyperspectral Image
Guansong Pang
Guansong Pang
Assistant Professor of Computer Science, Singapore Management University
Machine LearningData MiningComputer VisionAnomaly DetectionOpen-world Learning