VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对视觉和文本漂移问题,建立了首个跨域遥感图像分割基准VPRef,并通过低秩适应方法提高模型在不同领域的适应性。
📝 Abstract
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe performance degradation under a coupled dual-drift paradigm: visual domain drift from cross-spatial-resolution mismatches and spectral variations, alongside textual logic drift from unconstrained, variable user-input granularities. To mitigate these bottlenecks, this paper establishes the first cross-domain RRSIS benchmark, designated as the Vaihingen-Potsdam Referring (VPRef) dataset, comprising 46,972 language-image-annotation triplets organized into a three-tier linguistic hierarchy. Building upon this benchmark, we develop a tailored parameter-efficient domain adaptation baseline anchored on the Segment Anything Model (SAM3) via Low-Rank Adaptation (LoRA). Our framework counteracts visual distribution discrepancies through pseudo-label-driven self-training and addresses textual logic drift via random multi-granularity text prompt mixing. Crucially, the distribution of empirical metrics across ablative variants suggests a potential decoupling between cross-modal semantic robustification and visual domain alignment, demonstrating that linguistic variance drives fine-grained semantic invariance while pseudo-label propagation governs macro-scale spatial grid alignment. Extensive benchmarks demonstrate the proposed framework achieves superior cross-domain segmentation boundaries while modifying merely 1.08\% of the foundational parameter footprint, establishing a robust baseline for future multi-modal remote sensing domain adaptation research. The dataset and code will be available at https://github.com/quanweiliu/VPRef.
Problem

Research questions and friction points this paper is trying to address.

Referring Remote Sensing Image Segmentation
visual domain drift
textual logic drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-domain benchmark
Low-Rank Adaptation (LoRA)
pseudo-label-driven self-training
multi-granularity text prompt mixing
🔎 Similar Papers
No similar papers found.
Q
Quanwei Liu
College of Science and Engineering, James Cook University, Cairns, 4878, Australia
T
Tao Huang
College of Science and Engineering, James Cook University, Cairns QLD 4878, Australia and Center for AI and Data Science Innovation, James Cook University, Cairns QLD 4878, Australia
J
Jiaqi Yang
Department of Forest and Wildlife Ecology, University of Wisconsin-Madison, Madison, WI 53706 USA
Wei Xiang
Wei Xiang
Distinguished Professor, Cisco Research Chair of AI and IoT, La Trobe University
Internet of ThingsMachine LearningWireless Sensor NetworksWireless CommunicationsComputer