GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态无人机感知中因视差、平台运动和镜头畸变导致的空间对应问题,提出GAAT模型,通过几何感知的对齐方法增强跨模态融合。
📝 Abstract
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.
Problem

Research questions and friction points this paper is trying to address.

UAV Multimodal Perception
Spatial Correspondence
Cross-Modal Fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Aware Alignment
syncPATC
MG-Sparse-MMA
RA-QCGCL
UAVMeta and StateBench
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jingpu Yang
Beihang University, Beijing, China.
D
Debin Tang
Northeastern University, Shenyang, China.
Y
Yilin Sun
Beihang University, Beijing, China.
Fengxian Ji
Fengxian Ji
Northeast University
agent、Machine learnin、CV
J
Jiahua Zhu
Beihang University, Beijing, China.
Wenrui Ding
Wenrui Ding
Professor, Beihang University
Y
Yufeng Wang
Beihang University, Beijing, China.