Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a key limitation in existing backdoor attacks against vision-language models, which typically rely on static bindings and cannot flexibly specify arbitrary target semantics during inference. To overcome this, the authors propose a programmable backdoor mechanism that enables dynamic control over any desired output caption at inference time through a single poisoning step. The core innovation lies in an “arbitrary-to-arbitrary” captioning paradigm that decouples target selection from the poisoning process, eliminating the need for retraining. This is achieved by integrating a heuristic poisoning strategy with feature-space trigger steganography—specifically, norm-constrained perturbations and non-semantic image patch generation. The method maintains the model’s normal performance while achieving high attack success rates across arbitrary targets and effectively evading multiple state-of-the-art backdoor defenses.
📝 Abstract
Existing vision-language model (VLM) backdoors are usually treated as static vulnerabilities: one-to-one and N-to-N attacks bind one or more triggers to a finite set of targets before victim training. This assumption substantially underestimates the threat. We show that a single poisoning phase can implant a programmable backdoor into a VLM, allowing an attacker to choose previously unseen target-caption semantics at inference time and synthesize corresponding stealthy triggers on demand. Unlike fixed-mapping attacks, the proposed any-to-any caption-control paradigm decouples post-training target selection from poisoning, enabling dynamic control of target captions without retraining the VLM. Our method has two components. First, a heuristic poisoning strategy exposes the model to diverse trigger-caption pairs, encouraging it to learn a general trigger-as-instruction rule rather than memorize a specific backdoor pattern. Second, a feature-space trigger steganography method maps any attacker-specified target caption to a stealthy visual trigger, implemented as either a norm-controlled perturbation or a non-semantic patch. Once inserted into arbitrary images, these triggers cause the poisoned VLM to generate outputs semantically aligned with the chosen target caption, even when the target was unseen during poisoning. Extensive experiments show that our attack achieves high any-to-any caption-control success rates, preserves clean model utility, and remains effective under several classical backdoor defenses.
Problem

Research questions and friction points this paper is trying to address.

backdoor
vision-language models
programmable attack
trigger synthesis
any-to-any control
Innovation

Methods, ideas, or system contributions that make the work stand out.

programmable backdoor
vision-language models
any-to-any caption control
trigger steganography
poisoning attack
T
Tao Lin
Key Laboratory of System Software (Chinese Academy of Sciences), Beijing, China; Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
Gaojie Jin
Gaojie Jin
Lecturer (Assistant Professor), University of Exeter
Machine LearningStatistical LearningTrustworthy AIHuman-GenAI-Alignment
Z
Zongxin Liu
Key Laboratory of System Software (Chinese Academy of Sciences), Beijing, China; Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
P
Peng Wu
Key Laboratory of System Software (Chinese Academy of Sciences), Beijing, China; Institute of Software, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China
L
Lijia Yu
Institute of AI for Industries, Chinese Academy of Sciences, Nanjing, China