LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the feature manifold distortion and loss of semantic details in domain-generalized semantic segmentation caused by style randomization and feature normalization. To mitigate these issues, the authors propose the LASA framework, which uniquely integrates vision-language model priors with structural anchors from source domains to enable fine-grained text-to-source guided style transfer (TSGST). Furthermore, LASA introduces a domain-aware query adapter (DAQA) and a domain-aware decoder optimizer (DADO) that jointly recover discriminative semantic details suppressed during adaptation. Extensive experiments demonstrate that LASA significantly outperforms existing methods across multiple challenging benchmarks, effectively enhancing segmentation performance on unseen target domains.
📝 Abstract
Domain Generalization Semantic Segmentation (DGSS) focuses on generalizing knowledge from labeled source domains to unseen target domains where data is unavailable during the training phase. While conventional methods utilize style randomization or feature normalization to mitigate domain shifts, they often impair feature integrity. Specifically, style randomization distorts the underlying feature manifold due to its coarse-grained nature, while feature normalization suppresses discriminative, domain-sensitive semantic details owing to its rigid design. To address these limitations, we propose the Language-and-Source-Anchored Alignment (LASA) framework, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO). Concretely, the TSGST module addresses manifold distortion by utilizing source features as structural anchors and vision-language model (VLM) priors as fine-grained guidance. To restore suppressed discriminative and domain-sensitive details, the DAQA module recalibrates object queries via categorical guidance and domain-aware signatures, while the DADO module aligns the resulting query distributions with a shared classifier to ensure consistent categorical responses across domains. Extensive experiments on challenging benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches.
Problem

Research questions and friction points this paper is trying to address.

Domain Generalization
Semantic Segmentation
Domain Shift
Feature Integrity
Style Randomization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain Generalization
Semantic Segmentation
Vision-Language Model
Feature Alignment
Style Transfer
🔎 Similar Papers
No similar papers found.
J
Jinhong Zhu
Xiamen University, Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen, Fujian, China
W
Weiqi Yan
Xiamen University, Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen, Fujian, China
Shengchuan Zhang
Shengchuan Zhang
Xiamen University
computer visionmachine learning
L
Liujuan Cao
Xiamen University, Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen, Fujian, China