🤖 AI Summary
This study addresses the limitation of existing forecasting systems that rely predominantly on point predictions and thus fail to adequately characterize uncertainty for informed decision-making. To overcome this, the authors propose a hybrid framework that extends point forecasts from classical models—such as Theta, exponential smoothing, and ARIMA—into probabilistic forecasts by integrating error post-processing with model-specific, horizon-dependent uncertainty scaling. The approach calibrates forecast errors using historical simulation, conformal prediction, quantile regression, and GARCH-based methods, and systematically evaluates in-sample versus out-of-sample calibration performance. Empirical results on the M4 dataset demonstrate an average 4.6% reduction in Continuous Ranked Probability Score (CRPS). In-sample calibration consistently outperforms out-of-sample calibration, particularly over longer forecast horizons, thereby validating the effectiveness and practical utility of the proposed framework.
📝 Abstract
Many forecasting systems produce point forecasts even when decisions require information about uncertainty. We investigate whether post-processing methods can systematically improve upon traditional Gaussian predictive distributions constructed from in-sample residuals. We propose a hybrid framework that combines forecast error post-processing with model-specific scaling of forecast uncertainty across horizons. For a comprehensive evaluation, we apply historical simulation, conformal prediction, quantile regression, and GARCH-based post-processing to point forecasts generated by Theta, exponential smoothing, and ARIMA models. Using 14,407 monthly series from the M4 competition and forecast horizons of 1 to 12 months, we evaluate performance using the continuous ranked probability score and rank-based statistical comparisons. Averaged across horizons, all post-processing variants improve upon the benchmark predictive distributions, with gains of up to 4.6%. In-sample calibration outperforms its out-of-sample counterpart in 11 of the 12 model-method combinations, although the preferred post-processing method depends on the base model and forecast horizon. The advantage of in-sample calibration generally increases at longer horizons. Our results show that organisations can extend existing point-forecasting systems to provide useful uncertainty quantification without computationally intensive repeated model re-estimation.