Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过结合冻结的基础模型和两次视觉-语言模型查询,提高了零样本航空图像分割的准确性,解决了基础模型直接应用时遗漏重要特征的问题。
📝 Abstract
Global welfare often depends on the correct interpretation of aerial and satellite imagery. Acting on such imagery (mapping flooded ground, crop extent, or damaged infrastructure) demands pixel-level segmentation to ensure perfect class localization. Pretrained general foundation models, when applied directly, often miss important features and cannot always find all the classes belonging to a given scene, overlooking smaller objects that matter most. We use a single consumer-grade GPU running a vision-language model (VLM) to supply this missing guidance, improving segmentation while producing structured, auditable evidence that drives the result and can be inspected on its own. We fuse three approaches: the frozen foundation model that labels every pixel, and two queries to a VLM, one to choose the classes that matter, and one to locate the small objects the base model misses. Evaluating across four aerial datasets, we see consistent gains at each stage where the base model is competent.
Problem

Research questions and friction points this paper is trying to address.

aerial segmentation
zero-shot
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
zero-shot aerial segmentation
inference-time guidance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Teresa DiMeola
Dept. of Computer and Information Science, University of Mississippi
Charles Walter
Charles Walter
Assistant Professor of Computer Science, The University of Mississippi
Fog ComputingMobile and Wearable SecurityBluetooth SecurityComputer Science EducationSelf-Adaptive Systems
H
Hong Xiao
Dept. of Computer and Information Science, University of Mississippi