Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出NSPER及NSPER+R方法,通过结合新颖性和惊奇度来优化经验回放和探索过程,提高基于图像的强化学习效率。
📝 Abstract
Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from high-dimensional visual inputs. Traditional sampling often relies on random or suboptimal experience selection, leading to redundant updates and slow learning. Improving efficiency requires mechanisms that prioritize informative experiences while also encouraging effective exploration. Prioritized Experience Replay (PER) addresses part of this challenge by reusing high-value transitions, while intrinsic rewards promote the exploration of novel or uncertain states. However, their integration has not been extensively studied. This paper introduces Novelty and Surprise Prioritized Experience Replay (NSPER), which uses novelty to capture underrepresented states and surprise to expose gaps in the agent's understanding of the environment. We further extend this with NSPER+R, integrating these signals as intrinsic rewards to jointly improve replay quality and exploration. Experiments on DeepMind Control Suite tasks show that NSPER and NSPER+R improve training efficiency and convergence speed compared to existing methods in image-based RL.
Problem

Research questions and friction points this paper is trying to address.

Sample Efficiency
Image-based Reinforcement Learning
Experience Prioritization
Exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Novelty and Surprise
Prioritized Experience Replay
Intrinsic Rewards
🔎 Similar Papers
No similar papers found.
H
Hoda Yamani
Robot Learning Team and CARES Robotics Lab, Department of Electrical, Computer, and Software Engineering, University of Auckland, Auckland, New Zealand
Henry Williams
Henry Williams
University of Auckland
RoboticsMachine Learning
B
Bruce A. MacDonald
Robot Learning Team and CARES Robotics Lab, Department of Electrical, Computer, and Software Engineering, University of Auckland, Auckland, New Zealand