Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prevalent issues of coarse, misaligned, or incomplete manual annotations in remote sensing semantic segmentation, which often distort model evaluation. To tackle this, the authors propose a training-free, reference-free mask fidelity assessment method that constructs counterfactual image pairs—preserving and erasing the region within the mask—and leverages a frozen vision-language model to evaluate whether class-specific evidence is concentrated inside the mask and absent outside it. This approach enables, for the first time, reference-free auditing of annotation quality in remote sensing segmentation, revealing systematic labeling biases across categories and facilitating automatic refinement of supervision signals. The proposed Contrastive Mask Fidelity (CMF) metric achieves 81% agreement with expert judgments across ten remote sensing datasets, substantially outperforming existing methods, and CMF-guided supervision significantly enhances cross-domain transfer performance.
📝 Abstract
Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.
Problem

Research questions and friction points this paper is trying to address.

ground-truth masks
remote sensing
semantic segmentation
annotation quality
evaluation paradox
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Mask Fidelity
reference-free evaluation
semantic segmentation
remote sensing
vision-language model
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Shuaishuai Cao
Central South University
S
Shuwei Peng
University of Chinese Academy of Sciences
Meng Tang
Meng Tang
Assistant Professor, University of California, Merced
computer visionmachine learningoptimization
M
Min Huang
Jiangxi Normal University
Y
Youjin Wang
Renmin University of China
J
Jie Chen
Central South University
J
Jing Ouyang
Jiangxi Normal University
Zhiwei Zhai
Zhiwei Zhai
BGI research
Medical image analysismachine learning