NOVA: Normal-Side Modeling for Training-Free Zero-Shot Video Anomaly Detection

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对零样本视频异常检测中决策边界模糊和模态差异问题,提出NOVA框架,通过改进提示构建与视觉正常锚点增强正常侧模型。
📝 Abstract
Training-free zero-shot video anomaly detection (ZS-VAD) leverages vision-language models (VLMs) to localize anomaly instances from a predefined anomaly vocabulary, without providing any video. Existing CLIP-based methods often emphasize anomaly-side semantics, while the competing normality side remains less carefully formulated. We identify two key limitations in existing solutions: (i) blurred decision boundary: normal prompts may contain ambiguous verbs, such as running, that are semantically close to anomalies, reducing normal and abnormal separation in the VLM embedding space; and (ii) modality gap: poor alignment between features of textual normal anchors and visual frames. We propose NOVA, a training-free ZS-VAD framework that strengthens the normal side at both linguistic and visual levels. NOVA introduces Normality-Aware Prompt Construction (NA), which excludes anomaly-adjacent verbs and biases normal descriptions toward static, low-motion scenes. To overcome the text-vision modality gap, NOVA constructs a Visual Normality Anchor (VNA), which creates a weighted visual normal anchor from the initial frames of each test video, providing a video-specific normal reference without task-specific training or annotations. NOVA achieves 89.86 percent AUC on UCF-Crime and 95.07 percent AUC and 84.82 percent AP on XD-Violence, reaching state-of-the-art performance among comparable training-free zero-shot methods.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Video Anomaly Detection
Decision Boundary
Modality Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Normal-Side Modeling
Zero-Shot Video Anomaly Detection
Normality-Aware Prompt Construction
Visual Normality Anchor
W
Wei-Chih Yin
Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
Y
Yun-Ching Kao
Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
C
Cheng-Kuan Lin
Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan
Yu-Chee Tseng
Yu-Chee Tseng
College of AI, National Yang Ming Chiao Tung University
mobile computingwireless networkartificial intelligence