Bridging Modalities and Tasks: A Unified Hierarchical ViT for SAR-to-Optical Translation and Semantic Segmentation

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出BMT框架,通过共享的分层视觉变换器联合优化SAR到光学图像转换和语义分割,以提高图像解释性和下游任务性能。
📝 Abstract
Synthetic Aperture Radar (SAR) images have all-weather, day-and-night observation capabilities. However, compared with optical images, their speckle noise and non-intuitive scattering mechanism limit the interpretability of the images. Generative models for SAR-to-optical (S2O) conversion can improve visual interpretability, but existing methods often ignore the constraints on semantic structure, which are necessary for downstream tasks, for the sake of visual effects. We propose a unified collaborative dual-task learning framework, termed BMT (Bridging Modalities and Tasks), that jointly optimizes S2O image translation and semantic segmentation through a shared hierarchical Vision Transformer. The framework integrates: (1) a LocalViTBlock that fuses global self-attention with spatial depthwise convolution through a learnable gating mechanism; (2) an enhanced output module combining multi-scale refinement processing, color correction and anti-aliasing, which calibrates channel-level color statistics through feature fusion; (3) a ControlNet-style conditional injection mechanism that encodes SAR wavelet features and segmentation labels into a multi-scale feature pyramid and injects them at each encoder layer through zero-initialized convolution; (4) a bounded Kendall uncertainty weighting scheme that prevents either task from dominating the shared representation. We evaluate the framework under both paired and unpaired translation settings, on the public WHU-OPT-SAR paired dataset and a self-constructed unpaired ship dataset built from HRSID and DIOR, respectively. The experimental results show that the proposed method achieves competitive S2O translation quality and semantic segmentation performance. The dataset and source code have been publicly released at https://github.com/Lewisyuaner/BMT-S2O-main.
Problem

Research questions and friction points this paper is trying to address.

SAR-to-optical translation
semantic segmentation
visual interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Collaborative Dual-Task Learning
Hierarchical Vision Transformer
LocalViTBlock
ControlNet-style Conditional Injection
Bounded Kendall Uncertainty Weighting
🔎 Similar Papers
No similar papers found.
S
Siyuan Liu
School of Automation, Northwestern Polytechnical University, Xi’an 710129, China
X
Xuze Zhang
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710129, China
Y
Yongshun Wang
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710129, China
L
Licong Pan
School of Automation, Northwestern Polytechnical University, Xi’an 710129, China
H
Hang Liu
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710129, China
Huihui Li
Huihui Li
EE Ph.D, University of Central Florida
Digital CommunicaationGeneral Purpose Representation amd Association Machine