🤖 AI Summary
This study addresses the scarcity of high-frequency, low-cost sub-municipal income data in middle-income countries between censuses, which hampers effective social policy design. It presents the first systematic validation of Google Maps points of interest (POIs) as a high-frequency proxy for household income at the census tract level across São Paulo, covering 26,625 areas. The authors reduce the dimensionality of sparse POI features using principal component analysis (PCA) and non-negative matrix factorization (NMF), then integrate these with gradient boosting regression models. Employing spatially aware cross-validation to prevent data leakage, the optimal NMF-enhanced gradient boosting model achieves an R² of 0.65, demonstrating robust predictive performance. The analysis further identifies specific POI categories significantly associated with income, offering a novel paradigm for fine-grained socioeconomic monitoring.
📝 Abstract
Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade. We test whether the composition of crowd-sourced Google Maps Points of Interest (POIs) can serve as a high-frequency, low-cost proxy for household income across the 26,625 census sectors of the municipality of Sao Paulo. Using a theoretically motivated set of POI categories retrieved from Google Places, we represent each sector by its POI counts, decompose these high-dimensional, sparse features with principal component analysis (PCA) and non-negative matrix factorization (NMF), and train a sweep of regression models to predict census-derived income. Under a data leakage-aware spatial validation design the best model (NMF with gradient boosting) attains a held-out R^2 of 0.65, with performance stable across feature-extraction methods. Interpretable decompositions reveal which POI types carry the income signal. These results suggest that commercial, crowd-sourced geospatial data can complement conventional income statistics during intercensal periods, and we discuss extensions toward multidimensional poverty and the capabilities framework.