Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Multimodal large language models (MLLMs) often inadvertently memorize sensitive information—such as personally identifiable data or harmful content—during training, and multimodal prompts can be exploited by adversaries to extract such knowledge. Yet, systematic evaluation of multimodal forgetting remains absent. Method: We introduce UnLOK-VQA, the first benchmark for targeted sensitive-knowledge forgetting in MLLMs, built upon high-quality, human-curated image–text pairs. We design both white-box and black-box multimodal extraction attacks and propose a white-box forgetting method grounded in hidden-state interpretability. Contribution/Results: We find that erasing answer-related hidden-state representations is the most effective defense, and model scale positively correlates with post-forgetting robustness. Experiments demonstrate that multimodal attacks significantly outperform unimodal ones, establishing a foundational framework for secure forgetting research in MLLMs.