Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对多模态大语言模型中的后门问题,提出RACER框架,通过识别和修复层间不一致性异常来消除模型中的后门。
📝 Abstract
MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for conventional classifiers, show limited effectiveness on MLLMs, while MLLM-specific defenses mainly operate at inference time, filtering suspicious inputs without removing the backdoor embedded in the model. To address this gap and eliminate latent backdoors from MLLMs at their source, we present RACER, a model-level repair framework motivated by a key observation: backdoors induce abnormal layer-to-layer evolution in internal representations, which we term the layer-wise inconsistency anomaly. Importantly, this anomaly is modality-dependent, concentrating primarily in the token region encoding the trigger features that the backdoor model actually relies on. RACER therefore decomposes the fused representation into visual and textual token regions, normalizes their layer-wise inconsistency separately, and recomposes them using modality-aware weights over a deep-layer window, yielding a region-aware inconsistency objective that better captures localized backdoor-induced anomalies. Through a min-max optimization, this objective drives worst-case perturbation synthesis and adversarial fine-tuning against the resulting perturbation to repair the model, suppressing the deep representational directional shifts on which backdoor behaviors rely. RACER requires only 100 clean samples and no knowledge of the trigger, attack objective, or even whether the input model contains a backdoor. Evaluations on three open-source MLLMs across 36 backdoor settings spanning image, text, and multimodal triggers show that RACER reduces the average ASR to 1.1%, reaching 0% in 32 settings, while preserving clean-task utility on both backdoor and clean models.
Problem

Research questions and friction points this paper is trying to address.

MLLMs
backdoor risks
model-level repair
inference time defenses
layer-wise inconsistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Region-Aware Consistency
Layer-Wise Inconsistency Anomaly
Modality-Dependent Abnormality
Adversarial Fine-Tuning
Min-Max Optimization
💼 Related Jobs
No related jobs found.
Jiali Wei
Jiali Wei
Xi'an Jiaotong University
AI TestingAI Security
M
Ming Fan
Xi'an Jiaotong University
M
Mingkun Zhang
Xi'an Jiaotong University
Haoyu Wang
Haoyu Wang
Singapore Management University
AI Safety and SecurityCompiler Testing
J
Jun Sun
Singapore Management University
Guoheng Sun
Guoheng Sun
University of Maryland, College Park
Deep LearningNatural Language ProcessingMobile Computing
X
Xiaoning Ren
Xi'an Jiaotong University
H
Haijun Wang
Xi'an Jiaotong University
T
Ting Liu
Xi'an Jiaotong University