Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models
This work addresses the challenges of illumination variation, camera jitter, and scale changes in non-stationary aerial videos by proposing a robust real-time human detection method that integrates optical flow with deep features. The study introduces a novel framework that, for the first time, combines hierarchical extreme learning machines (H-ELM) with optical flow to achieve efficient detection. It systematically evaluates the trade-offs between accuracy and speed among supervised CNNs, pre-trained CNN feature extractors, and H-ELM in complex aerial scenarios. Experimental results on the UCF-ARG dataset show that the pre-trained CNN achieves an average accuracy of 98.09%, while H-ELM attains 95.9% accuracy with only 445 seconds of training time on a CPU, significantly outperforming conventional methods such as SVM and offering an effective balance between high accuracy and low computational cost.