Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究针对数据驱动政策可能对部分个体造成伤害的问题,提出了一种考虑反事实伤害的优化方法,通过优先级排序和跳跃式调整来改进政策设计。
📝 Abstract
Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of "first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.
Problem

Research questions and friction points this paper is trying to address.

policy intervention
individual harm
welfare improvement
counterfactual optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual optimization
lexical leapfrogging
harm-aware policy
🔎 Similar Papers