Policy choice in time series by empirical welfare maximization
Dynamic multivariate time series pose challenges including time-varying environments, historical dependence, dynamic causal effects, and statistical dependencies. Method: This paper proposes Time-series Empirical Welfare Maximization (T-EWM), the first extension of the empirical welfare maximization framework to dynamic time-series settings. T-EWM employs nonparametric potential outcome modeling and conditional welfare optimization to learn dynamic optimal policies. Contribution/Results: We establish theoretical guarantees, including conditional welfare consistency and a non-asymptotic upper bound on policy regret. In simulation studies and an empirical application to COVID-19 containment policy evaluation, T-EWM significantly improves policy welfare and achieves rapid regret convergence under limited samples. The framework provides a novel paradigm for dynamic decision-making that balances statistical rigor with practical feasibility.