🤖 AI Summary
This study addresses spatial heterogeneity in covariate-outcome relationships within small area estimation by proposing the Spatially Clustered Fay-Herriot (SC-FH) framework. Incorporating a Potts spatial cohesion term into a penalized likelihood formulation, the method employs sequential label updating and closed-form solutions to simultaneously estimate regression coefficients and achieve spatially consistent domain partitioning. Simulations demonstrate that SC-FH accurately recovers underlying mechanisms and enhances predictive accuracy. Furthermore, an empirical application to agricultural production in the Po River Valley successfully identifies two distinct farming regimes. The proposed approach improves upon direct estimators while effectively preventing over-shrinkage, thereby significantly enhancing the reliability of small area estimates in the presence of spatial non-stationarity.
📝 Abstract
Area-level small area estimation (SAE) models, such as the Fay--Herriot (FH) model, borrow strength across domains through covariates and random effects, but they can struggle when the relationship between the covariates and the outcome is spatially heterogeneous, that is, when it changes across the spatial domain of interest. We propose a spatially-clustered FH (SC-FH) framework that simultaneously (i) estimates cluster-specific regression coefficients and random effects variances and (ii) generates spatially coherent partitions of the geographical domain. Estimation maximizes a penalized likelihood that augments the FH likelihood with a Potts-type spatial cohesion term over the areal adjacency graph, through an efficient strategy that alternates between sequential label updates and closed-form FH updates within clusters. In simulation experiments run on the real geography of the application, the method recovers the latent regimes almost exactly whenever they are separated in the covariate--response space and improves prediction accuracy over the standard FH benchmark, with the spatial penalty acting as a stabilizer of both classification and estimation. An empirical application to the average standard output of farms in the Po Valley (Northern Italy) identifies two spatially compact production regimes with significantly different cluster-wise coefficients, and shows that the clusterwise predictor improves on the direct estimates while avoiding the over-shrinkage of the pooled model.