Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对废物分类中语义分割标记数据获取困难的问题,提出一种基于视觉-语言引导的伪标签方法,利用SAM和EVA-CLIP生成高质量伪标签以实现无监督领域自适应。
📝 Abstract
Obtaining labeled data for semantic segmentation in applied settings (e.g., autonomous driving, industrial waste sorting) is expensive and often infeasible at scale. We present a cross-modal pseudo-labeling pipeline that enables unsupervised domain adaptation without any target-domain annotations. The pipeline is built on two core foundation models: SAM generates class-agnostic region proposals, and EVA-CLIP assigns semantic labels based on region-text similarity, with confidence filtering ensuring that only reliable pseudo-labels are used for self-training a segmentation model. As an optional extension, BLIP provides language-grounded verification for ambiguous regions, thereby improving pseudo-label quality without altering the overall pipeline. Evaluated on two domain shifts, synthetic-to-real autonomous driving and, with a primary focus, lab-to-factory industrial waste sorting, the pipeline consistently improves over source-only baselines. Our results demonstrate that pseudo-label quality, not quantity, is a decisive factor in self-training under domain shift, and that cross-modal language grounding offers a practical path to reliable automatic annotation in deployment-critical applications.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Domain Adaptation
Semantic Segmentation
Pseudo-Labels
Waste Sorting
Cross-modal
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal pseudo-labeling
unsupervised domain adaptation
language grounding
self-training
semantic segmentation
🔎 Similar Papers