Technical Report on the CVPR 2026@AdvML Workshop Challenge
This work addresses the vulnerability of Vision-Language Agents (VLAs) in autonomous driving under multimodal adversarial attacks by organizing a challenge centered on DriveLM-style multi-view visual question answering. Participants are tasked with generating high-fidelity adversarial images and text perturbations with minimal textual distortion to mislead models into producing answers that deviate from reference responses, while evaluating transferability under both white-box and black-box settings. The study presents the first systematic assessment of multi-view multimodal attacks, introducing novel techniques such as QA-graph-guided budget allocation, feature-space optimization, and suffix-constrained textual perturbations. Key findings include the dominance of image-side attacks, the efficacy of scene-level optimization, and the susceptibility of layout-sensitive content, thereby establishing a benchmark for robustness evaluation and defense development in VLAs.