🤖 AI Summary
This study addresses the challenges of analyzing provincial poverty in Indonesia, where a small sample size (n = 34) and high-dimensional multicollinearity undermine the stability of conventional regression models. To tackle this, the authors develop a systematic comparative framework evaluating several regularization and machine learning approaches—including ridge regression, LASSO, elastic net, Bayesian shrinkage priors, spatial ICAR, and Bayesian additive regression trees (BART)—in terms of predictive performance and robustness. The results demonstrate that parametric linear shrinkage methods, particularly ridge regression, yield the most accurate and stable predictions, whereas more complex ensemble models tend to overfit. Notably, ICT skills emerge as a consistently significant negative predictor of poverty across all well-performing models, highlighting their potential as a strategic priority for development policy. This work offers a reliable modeling paradigm and empirical foundation for evidence-based policymaking in data-scarce settings.
📝 Abstract
Identifying the structural drivers of poverty in regional datasets is frequently hindered by small sample sizes and high multidimensional collinearity, which can result in unstable and misleading policy advice. This paper evaluates the provincial causes of poverty in Indonesia by addressing these specific statistical hazards. We employ a rigorous model-comparison framework designed for small samples ($n=34$) with high collinearity, comparing standard linear models with frequentist penalisation, Bayesian shrinkage priors, an adjusted spatial intrinsic conditionally autoregressive (ICAR) model, and complex machine learning ensembles. To ensure a robust evaluation, we measure predictive performance using strict Leave-One-Out Cross-Validation (LOOCV). The results demonstrate that algorithmic complexity is inherently risky in regional datasets: simple linear shrinkage models (Ridge, Elastic Net, LASSO) achieve the superior out-of-sample prediction, whereas complex ensembles like BART suffer from severe overfitting. Across all successful regularised models, ICT skills consistently emerge as the most stable proxy for lower provincial poverty. The primary contribution of this paper is demonstrating that, in data-constrained regional analysis, parametrically regularised linear shrinkage provides a more reliable mathematical foundation for isolating structural development priorities, such as ICT, than either naive OLS or unconstrained machine learning.