Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
提出了一种基于YOLO姿态估计和CLIP语义评分的轻量级两阶段框架,用于实时视频异常检测,无需光流或独立姿态估计器。
📝 Abstract
We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n-pose to detect persons and extract seventeen skeletal keypoints in a single forward pass. The second stage encodes each cropped person region through CLIP ViT-B/32 and computes cosine similarity against predefined textual descriptions of anomalous behaviors. This architecture eliminates the need for optical flow, standalone pose estimators, and density-based scoring modules. Experiments on CUHK Avenue, ShanghaiTech Campus, and a custom indoor dataset collected at Chulalongkorn University demonstrate an end-to-end throughput of approximately 51 FPS on an NVIDIA Titan XP GPU, a 3.36x speedup over the multi-feature baseline, while maintaining frame-level AUROC values of 89.26%, 70.26%, and 84.13%, respectively.
Problem

Research questions and friction points this paper is trying to address.

Real-Time
Video Anomaly Detection
YOLO
CLIP
Semantic Scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

YOLO v11n-pose
CLIP ViT-B/32
cosine similarity
real-time video anomaly detection
lightweight framework
💼 Related Jobs
No related jobs found.
V
Vanodhya G. Warnasooriya
Dept. of Electrical Engineering, Faculty of Engineering, Chulalongkorn University, Bangkok 10330, Thailand
A
Amir Hajian
Dept. of Electrical Engineering, Faculty of Engineering, Chulalongkorn University, Bangkok 10330, Thailand
W
Watchara Ruangsang
Media Technology Program, King Mongkut’s Univ. of Technology Thonburi, Bangkok 10150, Thailand
S
Supavadee Aramvith
Dept. of Electrical Engineering, Faculty of Engineering, Chulalongkorn University, Bangkok 10330, Thailand