Advantage-Driven Explicit Memory for Social Navigation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种具有非参数记忆的导航代理,通过记录关键事件前的步骤来减轻表示学习负担,并利用优势信号处理稀有高影响事件,提高对OOD情况的泛化能力。
📝 Abstract
Robot policies are predominantly learned with classical parametric variants of imitation learning or RL, where training stores the agent's behavior exclusively in the policy's network parameters, putting a heavy burden on the representation learning algorithm. We propose a new navigation agent equipped with non-parametric memory which explicitly indexes prior steps leading to critical events. The advantages are twofold: first, it allows the policy to outsource some of its behavior into an explicit memory; second, it encourages a form of continual learning by allowing an agent to collect data from its testing episodes during deployment and therefore to better generalize to OOD situations. In the context of social navigation, we show that this improves the agent's capability to retain sparse, high-cost failures, such as human collisions. If the policy is trained in simulation, this also naturally addresses the sim-to-real gap, partially, by basing some of the decision making on real data. We integrate the explicit memory into a recurrent PPO architecture and use hidden states for memory retrieval to capture continuous spatiotemporal dynamics. The goal of exploiting rare, high-impact events is achieved by leveraging the RL agent's advantage signals. We train our agent in simulation with a combination of photorealistic rendering and non-visual crowd simulation and show that the agent is robust with respect to OOD social behavior.
Problem

Research questions and friction points this paper is trying to address.

Social Navigation
Imitation Learning
Reinforcement Learning
Out-of-Distribution
Explicit Memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

non-parametric memory
continual learning
advantage-driven
social navigation
sim-to-real gap
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yeonsoo Park
Interdisciplinary Program in Artificial Intelligence, Seoul National University, Korea; the work was done while at Naver Labs Europe, France
Mattia Racca
Mattia Racca
Naver Labs Europe, France
G
Guillaume Bono
Naver Labs Europe, France
S
Steeven Janny
Naver Labs Europe, France
Gianluca Monaci
Gianluca Monaci
Naver Labs Europe
roboticscomputer visionmachine learning
Tomi Silander
Tomi Silander
Naver Labs Europe, France
Christian Wolf
Christian Wolf
Naver Labs Europe
AI for RoboticsMachine LearningComputer Vision