Efficient Adaptation For Remote Sensing Visual Grounding
Direct transfer of foundational vision-language models (e.g., Grounding DINO, OFA) to remote sensing visual grounding tasks suffers from significant performance degradation due to domain shift. Method: We propose a lightweight cross-domain adaptation framework. For the first time, we systematically investigate LoRA’s effectiveness across all modules of Grounding DINO; for OFA, we synergistically integrate BitFit and Adapter for parameter-efficient fine-tuning. Contribution/Results: Our method fine-tunes fewer than 10% of model parameters, reducing training cost by over 90% and substantially accelerating inference. It achieves state-of-the-art or competitive performance on multiple remote sensing visual grounding benchmarks. This work delivers a practical, low-overhead, high-performance, and deployment-friendly solution for multimodal remote sensing understanding.