Reinforcement Learning with Stochastic Reward Machines
Existing reward machines (RMs) assume noise-free reward signals, limiting their applicability in real-world settings characterized by sparse, action-sequence-dependent, and noisy rewards. Method: This paper proposes the Stochastic Reward Machine (SRM), a novel RM variant that explicitly models stochastic reward observations. We introduce constraint solving into RM learning for the first time, enabling automatic inference of state partitions and transition relations from agent exploration trajectories to synthesize a minimal SRM. Theoretical analysis establishes asymptotic convergence to an optimal policy under reward noise. Results: Experiments on two representative noisy-reward tasks demonstrate that our approach significantly outperforms existing RM-based methods and naive denoising baselines, validating its robustness and effectiveness in learning from unreliable reward signals.