Transformer-Based Dual-Optical Attention Fusion Crowd Head Point Counting and Localization Network

📅 2025-05-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address low accuracy in crowd counting and head localization under dense occlusion and low-light conditions in drone-captured scenes, this paper proposes an end-to-end point detection framework leveraging visible-infrared dual-modal inputs. The method employs a Transformer-based backbone integrated with dual-stream CNN feature extraction, cross-modal attention, and explicit feature alignment mechanisms. We introduce two novel modules: the Dual-Optical Attention Fusion Module (DAFP) and the Adaptive Dual-Modal Feature Decomposition Fusion Module (AFDF), which jointly mitigate systematic cross-modal misalignment. Additionally, a spatial random shift data augmentation strategy is adopted to enhance model robustness. Evaluated on the DroneRGBT and GAIIC2 benchmarks under dense, low-light scenarios, our approach achieves an 18.7% reduction in Mean Absolute Error (MAE) and a 23.4% improvement in head localization accuracy over state-of-the-art methods.

Technology Category

Application Category

📝 Abstract
In this paper, the dual-optical attention fusion crowd head point counting model (TAPNet) is proposed to address the problem of the difficulty of accurate counting in complex scenes such as crowd dense occlusion and low light in crowd counting tasks under UAV view. The model designs a dual-optical attention fusion module (DAFP) by introducing complementary information from infrared images to improve the accuracy and robustness of all-day crowd counting. In order to fully utilize different modal information and solve the problem of inaccurate localization caused by systematic misalignment between image pairs, this paper also proposes an adaptive two-optical feature decomposition fusion module (AFDF). In addition, we optimize the training strategy to improve the model robustness through spatial random offset data augmentation. Experiments on two challenging public datasets, DroneRGBT and GAIIC2, show that the proposed method outperforms existing techniques in terms of performance, especially in challenging dense low-light scenes. Code is available at https://github.com/zz-zik/TAPNet
Problem

Research questions and friction points this paper is trying to address.

Improves accuracy in dense crowd counting under UAV view
Enhances robustness for all-day counting using infrared fusion
Addresses misalignment issues in multi-modal image localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-optical attention fusion for crowd counting
Adaptive feature decomposition for precise localization
Spatial random offset for robust data augmentation
🔎 Similar Papers
No similar papers found.
Fei Zhou
Fei Zhou
HAUT
deep learningtarget detectionimage processing
Y
Yi Li
Neusoft Institute Guangdong, China
M
Mingqing Zhu
Airace Technology Co.,Ltd., China