AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

📅 2026-07-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of infrared remote sensing vision-language models (VLMs) to physically realizable adversarial attacks in safety-critical scenarios. The authors propose AirflowAttack, the first method to model thermal airflow turbulence as a prior for adversarial perturbations, leveraging a lightweight generator to synthesize input-agnostic, highly physically plausible perturbations that enable cross-model black-box attacks. By integrating CLIP surrogate model optimization, thermal airflow pattern regularization, and multi-task evaluation, the approach achieves an average attack success rate of 48.5% across five CLIP backbones, substantially outperforming existing infrared physical attacks. Furthermore, it induces up to a 38.2% relative drop in accuracy on six state-of-the-art VLMs, exposing a critical cognitive flaw wherein models misinterpret adversarial perturbations as genuine thermal signals.
📝 Abstract
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior. A lightweight generator synthesizes a single input-agnostic perturbation regularized toward physically plausible airflow patterns. Optimized on one surrogate CLIP model, it attains a mean zero-shot scene-classification attack success rate (ASR, the fraction of samples whose top-1 class changes) of 48.5% across five diverse CLIP backbones, far exceeding four IR-specific physical baselines (27.7--37.0%). Applied to six state-of-the-art VLMs, it cuts scene-classification accuracy by up to 38.2% relative, yet paradoxically makes some models more confident in their IR analysis, confabulating the perturbation as genuine thermal evidence such as temperature gradients and convection. Ablations show the airflow prior raises physical plausibility at no measurable cost to attack success. Together with a benchmark spanning eleven models and four tasks, these findings expose critical vulnerabilities in the rapidly expanding IR VLM ecosystem.
Problem

Research questions and friction points this paper is trying to address.

infrared remote sensing
vision-language models
adversarial robustness
thermal-airflow perturbations
physical plausibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

AirflowAttack
thermal-airflow perturbation
infrared remote sensing
vision-language models
adversarial robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Cong Su
Cong Su
Yale University
atomic engineeringelectron microscopy2D materialsquantum emitters
J
Jiaju Han
China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China
X
Xuemeng Sun
China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China
C
Chengyin Hu
China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China
Q
Qike Zhang
China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China
J
Jiujiang Guo
Tianjin University, Tianjin, China
Y
Yiwei Wei
China University of Petroleum-Beijing at Karamay, Karamay, Xinjiang, China
J
Jiahuan Long
Shanghai Jiao Tong University, Shanghai, China