🤖 AI Summary
Time-series cross-validation (TS-CV) exhibits systematic bias under finite-sample settings—even for well-specified models such as vector autoregressions (VAR) or those with martingale difference errors—due to inherent conflicts between temporal dependence and rolling or truncated validation strategies.
Method: Through rigorous theoretical analysis and mathematical derivation, the paper formally establishes, for the first time, that conventional TS-CV procedures—long presumed unbiased—are intrinsically biased. The analysis spans multiple mainstream time-series models and isolates the root cause: the incompatibility between serial dependence and fixed-window or horizon-limited validation protocols.
Contribution/Results: This work corrects a longstanding misconception in the literature regarding TS-CV’s validity, providing critical theoretical grounding for model selection in time-series forecasting. It issues a fundamental caution against uncritical application of standard CV in dynamic settings and motivates principled redesigns of validation frameworks tailored to dependent data.
📝 Abstract
It is well known that model selection via cross validation can be biased for time series models. However, many researchers have argued that this bias does not apply when using cross-validation with vector autoregressions (VAR) or with time series models whose errors follow a martingale-like structure. I show that even under these circumstances, performing cross-validation on time series data will still generate bias in general.