🤖 AI Summary
This study addresses the limitations of conventional forecast evaluation, which overemphasizes average accuracy while neglecting differences in risk and reliability relevant to real-world applications. Drawing an analogy between forecast loss differentials and financial return series, the authors introduce a suite of risk-adjusted metrics—including the Sharpe ratio, Sortino ratio, Omega ratio, maximum drawdown, and a newly proposed Edge Ratio—to systematically assess the robustness and informational value of predictive models. Integrating econometric models, machine learning algorithms, foundation models (TabPFN), and expert survey forecasts, the analysis reveals that although professional forecasters often exhibit lower average accuracy than machine learning models, they consistently outperform in risk-adjusted terms and achieve higher Edge Ratios. Certain machine learning approaches also demonstrate favorable risk characteristics for specific forecasting targets.
📝 Abstract
Average forecast accuracy is not the same as forecast reliability. I treat forecast loss differentials relative to a benchmark as a return series. I then evaluate these returns using risk-adjusted performance measures from finance, including the Sharpe ratio, Sortino ratio, Omega ratio, and drawdown-based metrics. I also introduce the Edge Ratio capturing a model's propensity to deliver uniquely informative predictions relative to the forecasting frontier. I apply this framework to U.S. macroeconomic forecasting, comparing econometric benchmarks, machine learning models, a foundation model (TabPFN), and the Survey of Professional Forecasters. While it is often feasible to beat professional forecasters in terms of average accuracy, it is much harder to beat them on a risk-adjusted basis. They rarely exhibit catastrophic failures and often achieve high Edge Ratios, plausibly reflecting the value of contextual judgment. Nonetheless, selected machine learning methods deliver attractive risk profiles for specific targets. The framework naturally extends to meta-analyses across targets, horizons, and samples, illustrated with a density forecast evaluation and the M4 competition.