Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了物体检测模型中的后门攻击问题,提出DistScan框架,通过监测预NMS预测分布变化来检测后门,无需额外训练或触发器知识。
📝 Abstract
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors when a hidden trigger is present. Existing detection methods either rely on trigger inversion or exploit architecture-specific assumptions, and critically, representative existing methods fail to generalize reliably to scene-level attacks, where a single trigger induces anomalous behavior across all objects in the scene simultaneously. We present DistScan, a backdoor detection framework based on a simple but previously unexploited observation: backdoor injection systematically shifts a model's pre-NMS prediction class distribution away from its training class frequencies, even on clean inputs without any trigger present. DistScan aggregates intermediate class predictions over a clean validation set and flags a model as backdoored if the resulting distribution deviates significantly from the training class frequencies, requiring no model weight access, no trigger knowledge, and no additional training. Extensive experiments on MS-COCO and PASCAL VOC across two architectures and three scene-level attack scenarios demonstrate that DistScan substantially outperforms existing methods, improving average detection accuracy over the best-performing applicable baseline by 27.32 percentage points.
Problem

Research questions and friction points this paper is trying to address.

backdoor attacks
object detection models
scene-level attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Backdoor Detection
Pre-NMS Prediction Distribution Shift
Scene-Level Attacks
Model Security