QuISE: Defense against Typographic Attacks on VLMs via Query-Irrelevant Semantic Editing

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of vision-language models (VLMs) to typographic attacks—where adversaries inject adversarial text to mislead model predictions—by proposing a training-free, model-agnostic black-box defense. The method identifies and replaces potentially harmful textual regions through semantically irrelevant edits and determines the final prediction based on the consistency of model outputs before and after editing. As the first general-purpose defense that requires neither model modification nor additional training, it is readily applicable to closed-source VLMs. Evaluated across three benchmarks, four attack configurations, and four mainstream VLMs, the approach substantially improves robustness, achieving average recovery rates of 67.9–75.0% with only 0.5–1.1% performance degradation on clean inputs.
📝 Abstract
Typographic attacks pose a critical threat to vision-language models (VLMs) by injecting misleading text into images and causing models to rely on adversarial textual cues rather than visual evidence. Existing defenses often require model-specific modifications, additional training, or access to internal model components, limiting their applicability to modern closed-source VLMs. In this paper, we propose QuISE, a model-agnostic, training-free black-box defense based on query-irrelevant semantic editing. QuISE first identifies text regions likely to affect the current query through influence-aware text localization. QuISE then replaces these regions with two semantically distinct replacement texts that are irrelevant to both the query and the image. The final answer is determined by answer consistency across the edited images. Extensive experiments on three typographic-attack benchmarks, four attack settings, and four VLMs show that QuISE consistently improves defended accuracy. QuISE achieves a recovery rate of 67.9-75.0% with a harm rate of 0.5-1.1%.
Problem

Research questions and friction points this paper is trying to address.

typographic attacks
vision-language models
adversarial defense
black-box defense
model-agnostic
Innovation

Methods, ideas, or system contributions that make the work stand out.

typographic attacks
vision-language models
model-agnostic defense
semantic editing
black-box defense
💼 Related Jobs
No related jobs found.
S
Shubin Lu
School of Software, Northwestern Polytechnical University, Xi’an, China
Jiaqi Yin
Jiaqi Yin
University of Maryland
EDALogic SynthesisFormal Verification
Y
Yihao Huang
Software Engineering Institute, East China Normal University, Shanghai, China