Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入Counterfactual Quotient Model解决强化学习中模型对与决策无关的高维现象过度建模的问题,直接从同步反事实展开中学习动作相关效应。
📝 Abstract
Reinforcement-learning models commonly predict complete future states, observations, or feature occupancies, even though action selection depends only on differences between the consequences of candidate actions. As a result, these models may devote substantial statistical and representational capacity to high-dimensional phenomena that evolve independently of the agent's current choice. We introduce the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions. Its canonical centered representation removes this common component while preserving every pairwise action comparison expressible by the modeled reward family. The implemented model learns these action-dependent effects directly from synchronized counterfactual rollouts, so shared stochastic dynamics cancel before function approximation rather than after complete futures have been predicted. We establish the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation. Controlled experiments in physics-based environments provide initial evidence for these properties: direct effect learning suppresses action-independent variation, supports previously unseen reward queries, and improves action ranking relative to models trained to predict absolute futures.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
action selection
future states
counterfactual
stochastic dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Quotient Model
action-dependent effects
synchronized counterfactual rollouts
decision sufficiency
common-mode invariance
🔎 Similar Papers
J
Junlin Chen
School of Computer Science and Engineering, Beihang University
R
Ruijie Wang
School of Computer Science and Engineering, Beihang University
Jianxin Li
Jianxin Li
School of Computer Science & Engineering, Beihang University
Big DataAIIntelligent Computing