Learning a Semantic Calibration Network for Open-Vocabulary Semantic Segmentation

πŸ“… 2026-06-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limited generalization to novel categories and semantic ambiguity in open-vocabulary semantic segmentation by proposing a Semantic Calibration Network (SCN). Built upon the CLIP pretrained model, SCN employs cross-attention to generate vision-aware pseudo-text embeddings and incorporates a category disambiguation module together with a dynamic logical fusion mechanism to model inter-class semantic relationships and refine the mask classification process. While preserving CLIP’s strong zero-shot generalization capability, SCN significantly enhances discriminative performance. It achieves state-of-the-art results across multiple mainstream benchmarks, substantially improving segmentation accuracy and robustness in open-vocabulary settings.
πŸ“ Abstract
Semantic image segmentation assigns a predefined category label to each pixel, has achieved significant progress lately. Open-Vocabulary Segmentation (OVS) extends the segmentation task from a fixed set to an open set, enabling the identification and segmentation of novel concepts based on arbitrary text inputs, such as category names or descriptions. In this paper, we propose a novel Semantic Calibration Network (SCN) for open-vocabulary semantic segmentation. Different from prior approaches that focus on feature aggregation or simple fine-tuning of pre-trained models, SCN refines the mask classification process by explicitly modeling the semantic correlations between classes, aiming to enhance the model's discriminative power while effectively preserving the generalization abilities of the pre-trained CLIP model. Specifically, SCN comprises two core components: Class Disambiguation (CD) and Logits Fusion (LF). First, a cross-attention mechanism is utilized to transform the text embeddings into visually aware pseudo-text embeddings, in order to derive an enhanced similarity score that complements the original mask-text similarity score. Subsequently, the Class Disambiguation module captures implicit inter-class dependencies through a residual architecture to effectively resolve semantic ambiguities. Finally, the Logits Fusion module dynamically integrates multifaceted semantic evidence to ensure that the model achieves a robust semantic consensus while maintaining CLIP's inherent generalization capability. Comprehensive experimental results on mainstream benchmarks demonstrate that the proposed method achieves significant performance improvements compared to state-of-the-art algorithms.
Problem

Research questions and friction points this paper is trying to address.

Open-Vocabulary Semantic Segmentation
Semantic Ambiguity
Novel Concept Segmentation
Text-to-Image Alignment
Zero-shot Segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Calibration Network
Open-Vocabulary Segmentation
Class Disambiguation
Logits Fusion
CLIP
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Y
Yang Sun
College of Computer and Data Science, Fuzhou University, Fuzhou, China.
Tao Wang
Tao Wang
Minjiang University
Computer VisionMachine Learning
Anastasia Ioannou
Anastasia Ioannou
Lecturer at University of Glasgow
Renewable energyoffshore windstochastic modellingenergy economics
G
Ge Xu
School of Computer and Big Data, Minjiang University, Fuzhou, China.