CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决开放世界物体检测中的语义模糊和未知物体过度抑制问题,提出CODE框架,通过跨模态校准、不确定性引导增强及动态抑制方法提升检测性能。
📝 Abstract
Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Suppression), a unified inference-time framework with three complementary components. Cross-Modal Joint Confidence Calibration injects global visual prototypes to calibrate text-driven known-class predictions. Uncertainty-Guided Universal Objectness Enhancement measures classification hesitation from local visual responses to strengthen potential unknown objects. Dynamic Outlier Suppression via Confidence Margin replaces rigid suppression with a margin-aware adjustment that preserves ambiguous out-of-distribution instances. Experiments on the Real-World Detection benchmark demonstrate that, with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively.
Problem

Research questions and friction points this paper is trying to address.

Open World Object Detection
semantic ambiguity
unidirectional text-to-vision matching
outlier penalties
unknown objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Calibration
Dynamic Suppression
Open World Object Detection
Uncertainty-Guided Enhancement
Confidence Margin
💼 Related Jobs
No related jobs found.
H
Hao Xu
Beijing Institute of Technology
Z
Zhaoning Shi
Beijing Institute of Technology
H
Hehe Jin
Beijing Institute of Technology
Bo Ma
Bo Ma
Beijing Institute of Technology
Computer VisionImage and Video ProcessingPattern Recognition