Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video anomaly detection (VAD) methods struggle to simultaneously achieve precise temporal localization and interpretable semantic descriptions. To address this limitation, this work proposes a human-like reasoning paradigm—the “Glance-think-Scrutinize” (GtS) framework—which introduces, for the first time, a training-free VAD agent capable of tool invocation and self-correction. The approach leverages multimodal large language models and integrates video cropping, dense frame resampling, cold-start fine-tuning, and a joint-reward reinforcement learning mechanism. Experimental results demonstrate that GtS significantly outperforms current training-free baselines, achieving superior detection accuracy and joint semantic-temporal performance while maintaining computational efficiency. The study also introduces a new benchmark, VAGU-T, along with a novel evaluation metric, JeAUG, to better assess comprehensive VAD capabilities.
📝 Abstract
Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-what" dissociation: traditional DNN-based methods localize when anomalies occur but lack semantic understanding, whereas LLM-based methods explain what happens but neglect precise temporal grounding. We attribute this to the absence of a unified reasoning paradigm. Inspired by how humans inspect surveillance videos - glancing globally to form temporal hypotheses, scrutinizing suspicious segments, and thinking iteratively to correct errors - we study this global-to-local paradigm from two perspectives. We first propose Glance then Scrutinize (GtS), a training-free framework using static and dynamic textual guidance for coarse-to-fine anomaly grounding and understanding, balancing accuracy and speed. To break the ceiling imposed by frozen external modules, we further propose a tool-augmented agentic VAD method, where a multimodal large language model learns to invoke a video cropping tool, inspect densely resampled frames, and self-correct mislocalized hypotheses, via cold-start supervised fine-tuning followed by reinforcement learning with a joint answer-grounding reward. For training and evaluation, we extend our prior VAGU benchmark into VAGU-T (Video Anomaly Grounding, Understanding, and Thinking), comprising 7,567 real-world videos over 21 anomaly categories with human-validated grounding, explanations, QA pairs, and chain-of-thought tool-calling traces. We further introduce JeAUG, a metric jointly evaluating semantic interpretability and temporal precision. Experiments show that GtS substantially surpasses training-free baselines, while the agentic model delivers both higher accuracy and faster inference.
Problem

Research questions and friction points this paper is trying to address.

Video Anomaly Detection
temporal grounding
semantic understanding
when-what dissociation
anomaly localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic reasoning
training-free VAD
tool-augmented MLLM
temporal grounding
video anomaly detection
🔎 Similar Papers
No similar papers found.
S
Shibo Gao
School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China.
P
Peipei Yang
Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China.; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China.
Xu-Yao Zhang
Xu-Yao Zhang
Institute of Automation, Chinese Academy of Sciences
Pattern RecognitionMachine LearningOCR
L
Linlin Huang
School of Electronic and Information Engineering, Beijing Jiaotong University, Beijing 100044, China.