Cross-modal learning for SAR target recognition using optical vision foundation models

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过使用基于光学数据训练的视觉基础模型为SAR图像分类提供类别级监督,提出了一种跨模态EO到SAR原型对齐框架以解决SAR目标识别难题。
📝 Abstract
Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Aperture Radar
Automatic Target Recognition
Cross-modal Learning
Foundation Models
Domain Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal learning
prototype alignment
DINOv3 vision foundation model
SAR classification
EO imagery
💼 Related Jobs
No related jobs found.
L
Lucas Hirsch
School of Engineering, University of Edinburgh, United Kingdom
J
James R. Hopgood
School of Engineering, University of Edinburgh, United Kingdom
J
Javid Khan
Leonardo UK, United Kingdom
Yoann Altmann
Yoann Altmann
Professor at Heriot-Watt University
Statistical Signal processingBayesian inferenceComputational Imaging
Mike Davies
Mike Davies
Director, Neuromorphic Computing Lab, Intel