Accounting for Preferential Sampling Using a Constructed Covariate

📅 2026-07-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses inference bias in geostatistics arising from preferential sampling—where sampling locations are stochastically dependent on the underlying spatial process—by proposing a simple yet effective covariate construction method. The approach characterizes the dependence between sampling locations and the spatial variable through the average distance to nearest neighbors of observed points and incorporates this metric into standard geostatistical models. Without requiring complex model extensions, the method remains compatible with existing inference tools. Monte Carlo simulations and empirical analyses of two real-world datasets—Portuguese fishery landings and Galician lead contamination biomonitoring—demonstrate that the proposed approach substantially mitigates preferential sampling bias and enhances the reliability of spatial inference.
📝 Abstract
In geostatistics, it is commonly assumed that sampling locations are selected independently of the underlying spatial process. In practice, however, this assumption is frequently violated. In fisheries, for example, sampling sites are often chosen to maximize expected catches, creating a stochastic dependence between the abundance process and the sampling design. Such preferential sampling can introduce substantial bias and compromise statistical inference. This study investigates the use of constructed covariates, based on average distances from nearest neighbours observations, that are able to mitigate preferential sampling. The inclusion of such covariate in the geostatistical model might be able to account for the stochastic dependence of sampling locations on the spatial variable. If this inclusion sufficiently captures the dependence, conventional methods of inference may be applied without resorting to more complex models. The proposed methodology is evaluated through an extensive simulation study that explores a variety of sampling scenarios and spatial configurations. Additionally, we demonstrate the practical utility of the approach using two real-world datasets: one on fishery landings provided by the Instituto Português do Mar e da Atmosfera, and another concerning lead pollution biomonitoring in Galicia. Results show that incorporating the constructed covariate can substantially reduce the impact of preferential sampling, enabling reliable inference with standard geostatistical tools. We also discuss practical challenges, limitations, and paths for future methodological development.
Problem

Research questions and friction points this paper is trying to address.

preferential sampling
geostatistics
spatial process
sampling design
statistical inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

preferential sampling
constructed covariate
geostatistics
nearest neighbour distance
spatial inference
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
A
Andreia Monteiro
Centre of Mathematics of the University of Minho (CMA T), Braga, Portugal
I
Isabel Natário
Department of Mathematics of the Nova School of Science and Technology, Center for Mathematics and Applications (NOVA Math), NOVA University of Lisbon, Caparica, Portugal
I
Ivone Figueiredo
Portuguese Institute for Sea and Atmosphere (IPMA), Center of Statistics and its Applications (CEAUL), Lisbon, Portugal
P
Paula Simões
Military Academy Research Center - Military University Institute (CINAMIL), Engineering Superior Institute of Lisbon (ISEL), Polytechnic Institute of Lisbon, Centre for Mathematics and Applications (NOVA Math), NOVA University of Lisbon, Portugal