Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言模型在未见类别上对对抗样本的性能下降问题,提出ADAPT方法,通过双提示机制分离稳健特征与伪稳健特征,提高模型鲁棒性。
📝 Abstract
While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen classes, leading to a rapid degradation in performance against adversarial examples of unseen classes as training progresses. We empirically identify that this degradation stems from the tendency of the model to learn pseudo-robust features (i.e., non-generalizable shortcuts). To mitigate this, we propose ADAPT (Adversarial Disentangled Prompt Tuning), a robust prompt tuning framework following the philosophy of ``Learning What Not to Learn''. Specifically, ADAPT uses a dual-prompt mechanism with a target prompt and a pool of decoy prompts. During training, the decoy prompts are guided to entrap diverse pseudo-robust features, while the target prompt is constrained to be orthogonal to the decoys in the embedding space to learn robust features. By disentangling the robust features from the pseudo-robust features, ADAPT effectively prevents robust generalization overfitting. We further provide an analysis showing that the orthogonal loss bounds the effect of shifts in pseudo-robust features on unseen classes, yielding a testing error guarantee. Empirically, extensive experiments demonstrate that ADAPT substantially improves the robustness of the target prompt on unseen classes. The code is available at https://github.com/cheny02/ADAPT-ACMMM2026.
Problem

Research questions and friction points this paper is trying to address.

adversarial prompt tuning
robust generalization overfitting
pseudo-robust features
unseen classes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Disentangled Prompt Tuning
dual-prompt mechanism
pseudo-robust features
orthogonal loss
🔎 Similar Papers
No similar papers found.
Yang Chen
Yang Chen
Professor, College of Computer Science and Artificial Intelligence, Fudan University
Social ComputingComputer NetworksApplied Machine Learning
Z
Zhan Zhuang
City University of Hong Kong, Hong Kong, China
Y
Yanbin Wei
Hong Kong University of Science and Technology, Hong Kong, China
Z
Zebin Chen
Southern University of Science and Technology, Shenzhen, China
Hua Liu
Hua Liu
Shanghai Jiao Tong University
hydrodynamics
Y
Yu Zhang
Southern University of Science and Technology, Shenzhen, China