Vision-Language Models Meet Meteorology: Developing Models for Extreme Weather Events Detection with Heatmaps
Existing vision-language models (VLMs) exhibit color perception bias and imprecise spatial localization when interpreting meteorological heatmaps, leading to unreliable explanations for extreme weather event detection (EWED). To address this, we formulate EWED as a vision-language question answering (VQA) task and introduce three key contributions: (1) ClimateIQA—the first domain-specific VQA dataset for meteorology; (2) SPOT, a novel algorithm that enhances precise localization of heatmap color boundaries and critical regions; and (3) Climate-Zoo, a family of meteorology-specialized VLMs. Experiments demonstrate that our approach elevates EWED accuracy from 0% to over 90%, substantially outperforming general-purpose VLMs. All datasets, source code, and pretrained models are publicly released, establishing a reproducible benchmark and foundational infrastructure for AI-driven meteorology.