Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing training-free video anomaly detection methods, which rely on the semantic capabilities of pretrained models yet lack explicit criteria for distinguishing hazardous events from normal behaviors. To overcome this, we propose the Contrastive Event Adjudication for Video Anomaly Detection (CEAVAD) framework, which introduces—for the first time—a falsifiable hypothesis contrasting dangerous and benign events. During inference, CEAVAD dynamically constructs an anomaly boundary and leverages a pretrained vision-language model to perform evidence matching and contrastive reasoning, enabling both training-free anomaly detection and interpretable localization. The method achieves state-of-the-art performance under the training-free paradigm across three mainstream benchmarks.
📝 Abstract
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality. To this end, we propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing explanations and video evidence. Specifically, CEAVAD first uses public-safety knowledge to construct hazard-benign event contrasts, pairing each hazard mechanism with a generic normal account and a mechanism-specific benign counterpart. It then determines whether the target interval better supports a hazard explanation or its benign competitor, yielding a revisable contrastive boundary proposal for the target. Finally, CEAVAD adjudicates between the competing explanations to determine whether the hazard hypothesis survives the video evidence, supporting both temporally localized anomaly detection and evidence-grounded explanations. Experiments on three widely used VAD benchmarks demonstrate that CEAVAD achieves state-of-the-art performance under the training-free paradigm.
Problem

Research questions and friction points this paper is trying to address.

video anomaly detection
training-free
anomaly decision criterion
event adjudication
contrastive reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
contrastive event adjudication
video anomaly detection
explanatory boundary
hazard-benign contrast
🔎 Similar Papers
No similar papers found.
W
Wenti Yin
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
X
Xiang Wang
Alibaba Group
H
Huaxin Zhang
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Hanqing Wang
Hanqing Wang
HUST ➡ Shanghai AI lab ➡ HKUST(gz)
MLLMEmbodied AIWorld ModelVLA
H
Hongbo Shao
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
C
Changxin Gao
Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology
Nong Sang
Nong Sang
Huazhong University of Science and Technology
Computer Vision and Pattern Recognition