TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning
This work proposes a text-driven unsupervised framework for video anomaly detection (VAD) that circumvents the reliance on scarce and heterogeneous annotated abnormal visual data. By leveraging temporal textual descriptions generated by large language models (LLMs) as surrogate training signals—without requiring any real anomalous videos—the method achieves cross-modal alignment between text and video through a frozen CLIP encoder. To capture both short- and long-term temporal dynamics, the approach introduces an event-evolution causal attention module that models the logical progression of events over time. Evaluated on the XD-Violence and UCF-Crime benchmarks, the proposed method significantly outperforms existing one-class and unsupervised VAD approaches, thereby eliminating the dependency on domain-specific abnormal visual examples.