Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in open-world object detection where ambiguous boundaries between known and unknown categories lead to difficulties in distinguishing hard negatives, unknown objects, and background clutter. To tackle this, the authors propose MSPO, a lightweight semantic calibration framework that, without altering existing detector architectures or incremental learning protocols, uniquely integrates semantic evidence derived from expanded textual descriptions with probabilistic objectness (PROB). This enables semantic calibration without requiring future class names, thereby preventing task degradation into open-vocabulary classification. Leveraging a frozen CLIP text encoder, MSPO constructs a category semantic space into which decoder queries are projected, enabling joint calibration based on semantic support and objectness. Evaluated on M-OWODB and S-OWODB benchmarks, MSPO significantly outperforms the PROB baseline, achieving up to a 2.7-point mAP gain on PASCAL VOC while maintaining strong unknown recall and effectively mitigating early-stage unknown confusion.
📝 Abstract
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space. However, visual objectness alone cannot determine whether an object-like query corresponds to a hard known instance, an unseen-category object, or background clutter, resulting in an ambiguous known-unknown decision boundary. We propose MSPO, a lightweight semantic calibration framework that augments PROB with task-aware known-category language priors while preserving its detector architecture and incremental learning protocol. For each currently known category, MSPO constructs an extended text description covering category attributes, visual appearance, typical scenes, and functional usage, and encodes it using a frozen CLIP text encoder. Decoder query features are projected into the same semantic space to estimate their support from the current known-category semantics. This semantic evidence is fused with PROB's visual objectness to calibrate known and unknown predictions without turning OWOD into open-vocabulary classification. Importantly, MSPO never uses future-category names, and all unseen categories remain unnamed during evaluation. Experiments on M-OWODB and S-OWODB show that MSPO improves the strong PROB baseline on the main aggregate metrics while retaining competitive unknown recall. It also improves early unknown-confusion metrics and raises PASCAL VOC final mAP by up to 2.7 points. These results demonstrate that known-category language semantics provide an effective calibration signal for probabilistic objectness under the standard OWOD setting.
Problem

Research questions and friction points this paper is trying to address.

open-world object detection
unknown object discovery
semantic ambiguity
probabilistic objectness
known-unknown decision boundary
Innovation

Methods, ideas, or system contributions that make the work stand out.

open-world object detection
semantic calibration
probabilistic objectness
language priors
CLIP
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
W
Weijun Tian
School of Computer Science and Engineering, Beihang University
Rui Liu
Rui Liu
University of Science & Technology of China
solar physics