Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models

📅 2018-02-12
🏛️ Computational Intelligence and Neuroscience
📈 Citations: 42
Influential: 2
📄 PDF
🤖 AI Summary
This work addresses the challenges of illumination variation, camera jitter, and scale changes in non-stationary aerial videos by proposing a robust real-time human detection method that integrates optical flow with deep features. The study introduces a novel framework that, for the first time, combines hierarchical extreme learning machines (H-ELM) with optical flow to achieve efficient detection. It systematically evaluates the trade-offs between accuracy and speed among supervised CNNs, pre-trained CNN feature extractors, and H-ELM in complex aerial scenarios. Experimental results on the UCF-ARG dataset show that the pre-trained CNN achieves an average accuracy of 98.09%, while H-ELM attains 95.9% accuracy with only 445 seconds of training time on a CPU, significantly outperforming conventional methods such as SVM and offering an effective balance between high accuracy and low computational cost.

Technology Category

Application Category

📝 Abstract
Human detection in videos plays an important role in various real life applications. Most of traditional approaches depend on utilizing handcrafted features which are problem-dependent and optimal for specific tasks. Moreover, they are highly susceptible to dynamical events such as illumination changes, camera jitter, and variations in object sizes. On the other hand, the proposed feature learning approaches are cheaper and easier because highly abstract and discriminative features can be produced automatically without the need of expert knowledge. In this paper, we utilize automatic feature learning methods which combine optical flow and three different deep models (i.e., supervised convolutional neural network (S-CNN), pretrained CNN feature extractor, and hierarchical extreme learning machine) for human detection in videos captured using a nonstatic camera on an aerial platform with varying altitudes. The models are trained and tested on the publicly available and highly challenging UCF-ARG aerial dataset. The comparison between these models in terms of training, testing accuracy, and learning speed is analyzed. The performance evaluation considers five human actions (digging, waving, throwing, walking, and running). Experimental results demonstrated that the proposed methods are successful for human detection task. Pretrained CNN produces an average accuracy of 98.09%. S-CNN produces an average accuracy of 95.6% with soft-max and 91.7% with Support Vector Machines (SVM). H-ELM has an average accuracy of 95.9%. Using a normal Central Processing Unit (CPU), H-ELM's training time takes 445 seconds. Learning in S-CNN takes 770 seconds with a high performance Graphical Processing Unit (GPU).
Problem

Research questions and friction points this paper is trying to address.

human detection
aerial video
real-time
nonstatic camera
dynamic environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

automatic feature learning
aerial human detection
deep models
optical flow
real-time detection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Nouar Aldahoul
Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia
A
Aznul Qalid Md. Sabri
Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia
A
Ali Mohammed Mansoor
Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia