🤖 AI Summary
This study addresses inefficient labor allocation and absent quality control standards in crowdsourced annotation of multi-source remote sensing imagery by presenting the first empirical comparison of annotations across UAV, manned aircraft, and satellite data. Multi-level quality audits reveal that low-resolution imagery exhibits a revision rate of 36.95%, underscoring the limitations of uniform review strategies. Accordingly, this work proposes an adaptive auditing mechanism tailored for multi-source dataset construction, which effectively overcomes traditional quality control bottlenecks. The proposed framework provides both theoretical foundations and practical paradigms for optimizing resource allocation in crowdsourced annotation and enhancing overall data quality. These findings offer actionable guidance for improving annotation workflows in heterogeneous remote sensing applications, demonstrating significant potential for advancing reliable large-scale geospatial dataset development through differentiated quality assurance protocols.
📝 Abstract
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.