Structuring Value Representations via Geometric Coherence in Markov Decision Processes
This work addresses the challenges of instability and low sample efficiency in value function estimation within reinforcement learning by introducing, for the first time, an order-theoretic perspective. The authors formulate value learning as a partially ordered set (poset) learning problem and propose the GCR-RL framework, which progressively refines a hyper-poset structure guided by temporal difference signals to ensure geometric consistency in value representations. Building on this foundation, they develop two novel algorithms compatible with both Q-learning and Actor-Critic architectures, accompanied by theoretical convergence guarantees. Empirical evaluations demonstrate that the proposed approach significantly improves sample efficiency and training stability across a variety of tasks, outperforming several strong baselines.