A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection

📅 2025-03-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Open-vocabulary object detection (OVD) suffers from weak generalization to unseen categories, primarily due to coarse-grained alignment between detector features and the CLIP embedding space, hindering effective semantic knowledge transfer. To address this, we propose a three-level semantic distillation framework: (1) instance-level modeling of single-object visual relationships; (2) category-level, text-guided novel-class-aware classification; and (3) image-level multi-object contextual contrastive distillation—systematically transferring CLIP’s instance-, category-, and image-level generalizable semantics. This is the first framework enabling cross-granularity collaborative knowledge distillation without additional text annotations. On OV-COCO with a ResNet-50 backbone, our method achieves 46.4% AP on novel classes, significantly outperforming state-of-the-art methods. Ablation studies quantitatively validate the contribution of each component.

Technology Category

Application Category

📝 Abstract
Open-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, eg, CLIP, to inherit its generalizable recognition ability so that detectors can recognize new or novel objects. However, previous works directly align the feature space with CLIP and fail to learn the semantic knowledge effectively. In this work, we propose a hierarchical semantic distillation framework named HD-OVD to construct a comprehensive distillation process, which exploits generalizable knowledge from the CLIP model in three aspects. In the first hierarchy of HD-OVD, the detector learns fine-grained instance-wise semantics from the CLIP image encoder by modeling relations among single objects in the visual space. Besides, we introduce text space novel-class-aware classification to help the detector assimilate the highly generalizable class-wise semantics from the CLIP text encoder, representing the second hierarchy. Lastly, abundant image-wise semantics containing multi-object and their contexts are also distilled by an image-wise contrastive distillation. Benefiting from the elaborated semantic distillation in triple hierarchies, our HD-OVD inherits generalizable recognition ability from CLIP in instance, class, and image levels. Thus, we boost the novel AP on the OV-COCO dataset to 46.4% with a ResNet50 backbone, which outperforms others by a clear margin. We also conduct extensive ablation studies to analyze how each component works.
Problem

Research questions and friction points this paper is trying to address.

Detects objects beyond training annotations using vision-language models.
Improves semantic knowledge transfer from CLIP in object detection.
Enhances recognition of novel objects through hierarchical semantic distillation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical semantic distillation for object detection
Instance-wise, class-wise, image-wise semantic learning
Enhanced novel object recognition using CLIP model
🔎 Similar Papers
No similar papers found.
Shenghao Fu
Shenghao Fu
Sun Yat-sen University
computer visionobject detectionlarge multi-modal models
Junkai Yan
Junkai Yan
Insta360
Self-supervised learningObject detectionMultimodal learning
Qize Yang
Qize Yang
Tongyi Lab, Alibaba Group
Computer VisionDeep Learning
X
Xihan Wei
Tongyi Lab, Alibaba Group, China
X
Xiaohua Xie
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, Guangdong, China, also with the Guangdong Key Laboratory of Information Security Technology, Sun Yat-sen University, Guangzhou, Guangdong, China, and also with the Key Laboratory of Machine Intelligence and Advanced Computing, Sun Yat-sen University, Ministry of Education, Guangzhou, Guangdong, China
Wei-Shi Zheng
Wei-Shi Zheng
Professor @ SUN YAT-SEN UNIVERSITY
Computer VisionPattern RecognitionMachine Learning