ZODS-RS -- Zero-training Oriented Detection & Segmentation for Remote Sensing

📅 2026-06-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing training-free methods struggle to achieve unified detection and segmentation across viewpoints and platforms in remote sensing and UAV imagery, particularly under challenges such as oriented geometry, scale/rotation variations, and dense small objects. This work proposes ZODS-RS, the first training-free, closed-form unified framework that integrates DINOv3 dense features with SAM-style proposals to enable axis-aligned bounding box detection and instance segmentation. It introduces prototype purification (PP), rotation-scale equivariant matching (R-SEM), and uncertainty-aware pixel fusion (UAM) to address these challenges effectively. Evaluated on FAIR1M, xView, and a newly curated UAV dataset, ZODS-RS significantly outperforms baseline methods, achieving a 30.70 AP improvement on small objects over Grounded-SAM, thereby demonstrating robust performance in complex scenes and cross-domain generalization.
📝 Abstract
Remote-sensing and UAV applications need models that generalize across platforms and viewpoints without task-specific training. Yet training-free pipelines often falter on oriented geometry, scale/rotation variation, and crowded ports or airfields, and rarely unify detection and segmentation. We introduce ZODS-RS, a training-free, closed-form pipeline that outputs horizontal boxes (HBB) and instance masks. Built on DINOv3 dense features and SAM-style proposals, ZODS-RS chains: PP (prototype purification via Tyler covariance), R-SEM (rotation-scale equivariant matching with separable kernels and global Hungarian assignment), and UAM (uncertainty-aware pixelwise merging with adaptive priors and optional negative prototypes). A lightweight CWLA fuses multiple DINOv3 layers. On FAIR1M (HBB) we obtain $\mathrm{mAP}_{0.50:0.95}=\mathbf{13.06}$ and $\mathrm{AP}_S=\mathbf{2.93}$ \emph{(class-averaged over ship/airplane)}; on xView (HBB) we report $\mathrm{mAP}=\mathbf{16.69}$. On our UAV dataset, ZODS-RS achieves mask $\mathrm{mIoU}=\mathbf{31.10}$ and improves small-object AP by $\mathbf{+30.70}$ over Grounded-SAM on a single 5090. This work offers a unified, \emph{no-training} solution for horizontal-box detection plus instance segmentation in aerial imagery; provides explicit closed-form formulations for PP/R-SEM/UAM tightly coupled with DINOv3; and demonstrates \emph{consistent} gains on small and crowded targets and under cross-domain shifts while keeping deployment simple.
Problem

Research questions and friction points this paper is trying to address.

remote sensing
zero-training
oriented object detection
instance segmentation
cross-domain generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-training
oriented object detection
instance segmentation
rotation-scale equivariance
DINOv3
💼 Related Jobs
No related jobs found.
Z
Zuan Gu
Northeastern University, China
T
Tianhan Gao
Northeastern University, China
L
Langxu Zhao
Northeastern University, China