Stable Policy Learning

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了在基于证据的政策制定中,如何通过算法稳定性来平衡预期福利与抽样风险,并提出了一种称为政策投票袋装的方法以提高决策的稳健性。
📝 Abstract
In evidence-based policymaking, typically one experimental sample is observed, then a learned policy recommendation is implemented at scale. Policies learned from the experimental data can perform well in expected welfare, yet random sampling in the experiment can produce recommendations with poor welfare outcomes. In this paper, we ask: how should policy learning algorithms balance expected welfare against sampling risk? Our main contribution is to show that algorithmic stability plays a central role in characterizing and navigating the tradeoff. Intuitively, if a policy learning algorithm's recommendation remains stable when one experimental unit is replaced, then that algorithm has limited sampling risk. We propose a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities. Relative to using one subsample, averaging across subsamples preserves expected welfare and improves expected utility for a risk-averse researcher. We derive sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.
Problem

Research questions and friction points this paper is trying to address.

evidence-based policymaking
expected welfare
sampling risk
algorithmic stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

algorithmic stability
policy-vote bagging
sampling risk
🔎 Similar Papers
2024-07-09Neural Information Processing SystemsCitations: 3
H
Harvey Barnhard
Harvard University
G
Giacomo Opocher
University of Mannheim
Rahul Singh
Rahul Singh
Harvard University
econometricscausal inferencemachine learning