Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the often-overlooked interplay between exploration and memory in partially observable reinforcement learning. The authors propose a formal framework based on observation-anchored reward machines to disentangle structural from latent sparse rewards, enabling a systematic investigation of how episodic exploration bonuses interact with diverse neural memory architectures—such as LSTMs and Transformers—across three distinct memory acquisition settings. Their experiments reveal that reward structure, rather than density, primarily governs the exploration–memory interaction; identical exploration bonuses can amplify, balance, or obscure differences among memory architectures; and only dense rewards that directly supervise the required latent memory can neutralize exploration incentives. Furthermore, minor penalties often trap policies in suboptimal equilibria, an issue effectively mitigated by exploration bonuses. These findings demonstrate that exploration and memory serve complementary, not substitutive, roles.
📝 Abstract
In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interaction unmeasured, and standard notions of sparse reward conflate temporal signal density with what the reward actually supervises. We present a controlled study crossing episodic exploration bonuses with diverse neural memory architectures across three environments that vary how the content of memory is acquired. An identical bonus signal yields three distinct interaction patterns: it amplifies architectural capacity differences where memory content must be actively discovered and retained unsupervised; equalizes architectures to a shared ceiling where the content, once sought out, is a single reward-supervised cue; and is null where the observation stream is purely scheduled. Controlled reward manipulations verify that these patterns track reward structure rather than density: a dense reward neutralizes a bonus only if it directly supervises the required latent memory, and a small avoidable penalty on exploratory actions (leaving the optimum unchanged) induces policy convergence to suboptimal stationary states, which either bonus resolves. We then formalize reward sparsity with observation-anchored reward machines, separating structural sparsity (an automaton reproduces the return without the task-required history) from potential sparsity (the one-step reward misprices local exploratory actions); the resulting vocabulary organizes the three regimes by the retention burden each task exposes. Together, these results show exploration and memory are complements, not substitutes: a bonus induces exposure, and only memory converts exposure into return.
Problem

Research questions and friction points this paper is trying to address.

partially observable reinforcement learning
episodic exploration
neural memory
reward sparsity
memory-exploration interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

episodic exploration
neural memory
reward sparsity
reward machines
partially observable reinforcement learning
🔎 Similar Papers
No similar papers found.
J
Jai Malegaonkar
Department of Computer Science and Engineering, UC San Diego, United States
Rohan Patil
Rohan Patil
UC San Diego
Machine LearningData Science
H
Henrik I. Christensen
Department of Computer Science and Engineering, UC San Diego, United States