Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in open-vocabulary instance segmentation and open-set panoptic segmentation, including noisy pseudo-labels, weak vision–language alignment, and difficulties in handling out-of-vocabulary categories. To tackle these issues, the authors propose a multimodal pseudo-labeling and training framework that integrates pretrained models such as Grounded SAM, LLaVA, and CLIP. The approach employs a target-vocabulary-guided pseudo-labeling mechanism, CLIP-driven synonym filtering, and GPT-enhanced caption reconstruction to construct semantically consistent vision–text pairs. By jointly optimizing an extended visual grounding loss, a semantic consistency loss, and a generative caption reconstruction loss, the model achieves significantly improved generalization to unseen categories. Evaluated on the COCO benchmark, the method sets new state-of-the-art results in both tasks.
📝 Abstract
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, we propose a multimodal framework that leverages pre-trained vision-language models for automatic pseudo-label generation, CLIP-guided synonym filtering, and GPT-based caption reconstruction. In our target-vocabulary-assisted pseudo-labeling setting, the framework first constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP, providing multimodal supervision without manual annotation. We then enhance visual-textual alignment through three complementary training objectives: an extended grounding loss that incorporates visually grounded synonyms, a semantic consistency loss, and a generative caption reconstruction loss. Extensive experiments on the COCO dataset demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches under this protocol, achieving substantial improvements on both OVIS and OSPS benchmarks.
Problem

Research questions and friction points this paper is trying to address.

open-vocabulary instance segmentation
open-set panoptic segmentation
pseudo-labeling
visual-textual grounding
out-of-vocabulary
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal pseudo-labeling
open-vocabulary segmentation
visual-textual alignment
synonym filtering
caption reconstruction
🔎 Similar Papers
D
Duy Tran Thanh
Department of Electronic Engineering, Seoul National University of Science and Technology, 232 Gongneung-ro, Nowon-gu, Seoul, 01811, South Korea
Y
Yeejin Lee
Department of Electrical and Information Engineering, Seoul National University of Science and Technology, 232 Gongneung-ro, Nowon-gu, Seoul, 01811, South Korea
Byeongkeun Kang
Byeongkeun Kang
Chung-Ang University
Computer VisionArtificial IntelligenceRobotics