Breaking Habits: On the Role of the Advantage Function in Learning Causal State Representations
In reinforcement learning, agents often suffer from policy-induced spurious correlations—termed “policy confounding”—arising from feedback loops between policies and observations, leading to poor generalization on unseen trajectories. This work establishes, for the first time, that the advantage function not only reduces variance in policy gradient estimates but also decouples such spurious policy-induced dependencies, thereby facilitating causal state representation learning. Building upon the policy gradient framework, we propose advantage-based state representation normalization of action-value functions, integrating principles from causal inference with empirical validation. Theoretical analysis proves that our method mitigates policy confounding by decorrelating non-causal state-action associations. Experiments across diverse benchmark tasks demonstrate consistent improvements in out-of-distribution trajectory generalization. Our approach yields a novel, interpretable, and generalizable paradigm for robust reinforcement learning—grounded in causal reasoning and empirically validated.