A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对空中-地面跨视角指代人物检测问题,提出A-PAIR基准及身份一致的定位框架,通过因子化标注与参照对齐方法有效构建数据集,并利用身份一致的指代定位框架提高检测性能。
📝 Abstract
Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target before ground and aerial agents can coordinate downstream actions. Existing referring expression comprehension and open-vocabulary grounding methods do not jointly account for cross-view identity consistency, making them insufficient for Air-Ground Cross-View Referring Person Detection (AGCV-RPD), which involves similar pedestrian distractors, weak aerial appearance cues, and cross-view identity consistency. To study this problem, we introduce Air-Ground Paired Identity-Aware Referring (A-PAIR), the first comprehensive AGCV-RPD benchmark, containing 22,137 cross-view referring samples. To construct A-PAIR efficiently, we propose Factorized Annotation and Referential Alignment (FARA), a semi-automatic annotation framework that generates factorized referring descriptions and identity-consistency supervision at reduced cost. We propose Identity-Consistent Referring Grounding (ICRG), a framework that combines factorized referential grounding, candidate-completeness supervision, and cross-view consistency calibration for joint air-ground pair selection. ICRG improves ground, aerial, and pair-level detection over strong baselines, increasing pair F1 from 16.65% to 22.28%. These results show that AGCV-RPD requires paired detection and identity-consistent reasoning.
Problem

Research questions and friction points this paper is trying to address.

Air-Ground Cross-View Referring Person Detection
Cross-View Identity Consistency
Pedestrian Distractors
Aerial Appearance Cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Identity-Consistent Referring Grounding
Factorized Annotation and Referential Alignment
Air-Ground Cross-View Referring Person Detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhoupeng Guo
School of Automation, Southeast University, Nanjing, China
X
Xinjie Yao
Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China
Y
Yunqi Zhu
School of Computer Science and Engineering, University of New South Wales, Sydney, Australia
Z
Zhihe Fan
School of Sports Training, Tianjin University of Sport, Tianjin, China
S
Siqi Zhao
Tianjin University, Tianjin, China
J
Jianjun Chen
Tianjin University, Tianjin, China
Y
Yichen Dong
Tianjin University, Tianjin, China
Y
Yan Fan
National University of Defense Technology, Changsha, China
Pengfei Zhu
Pengfei Zhu
Professor, College of Intelligence and Computing , Tianjin University
computer visionpattern recognitionmachine learning