GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对无人机图像中视觉定位问题,提出GrabVG框架,通过预注意假设搜索和图注意力特征绑定两阶段方法,提高密集场景下目标定位准确性。
📝 Abstract
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distributed, and visually similar objects creates high visual redundancy, while repetitive local configurations give rise to strong topological ambiguity. Existing approaches mainly focus on visual--language feature alignment or dense contextual interaction, yet they struggle to distinguish subtle inter-instance differences and effectively exploit spatial topological structures, leading to inaccurate grounding in highly crowded scenarios. To address these challenges, we propose $\textbf{GrabVG}$, a novel visual grounding framework inspired by human visual search. GrabVG explicitly decomposes grounding into two sequential stages: $\textit{preattentive hypothesis search}$ and $\textit{graph-attentive feature binding}$. Specifically, we first generate a compact set of reliable object hypotheses through distillation-guided proposal induction and text-aware hypothesis filtering, substantially reducing background distractions and semantic mismatches. These hypotheses are then organized into a sparse graph, where language-guided intra-instance visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention, enabling efficient spatial reasoning and accurate target localization. Extensive experiments on AerialVG and AerialSense show that GrabVG achieves a favorable accuracy--speed trade-off, reaching 67.31$\%$ and 80.34$\%$ Acc@0.5 and outperforming the corresponding baselines by 10.55 and 8.76 percentage points, respectively.
Problem

Research questions and friction points this paper is trying to address.

Visual Grounding
UAV Imagery
Topological Ambiguity
Visual Redundancy
Spatial Topological Structures
Innovation

Methods, ideas, or system contributions that make the work stand out.

preattentive hypothesis search
graph-attentive feature binding
visual grounding
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Chaowei Wang
Chaowei Wang
Associate Professor of Beijing University of Posts & Telecommunications
wireless communications
Yan Di
Yan Di
Harbin Institute of Technology, Shenzhen
pose estimation
J
Jingjun Sun
Northwestern Polytechnical University
B
Baozhe Liu
The Hong Kong Polytechnic University
J
Jiaxu Tian
Northwestern Polytechnical University
Y
Yuheng Li
Northwestern Polytechnical University
G
Guangqian Guo
Northwestern Polytechnical University
S
Shan Gao
Northwestern Polytechnical University