Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing personalized federated reinforcement learning methods, which overly rely on extrinsic rewards and suffer from insufficient exploration in non-stationary or sparse-reward environments, leading to weak policy personalization and low sample efficiency. To overcome these challenges, we propose the first exploration-driven personalized federated reinforcement learning framework that integrates intrinsic motivation. Specifically, clients leverage Random Network Distillation (RND) to generate intrinsic rewards that enhance local exploration, while the server aggregates only minimal novelty summaries and broadcasts a global exploration prior to coordinate diverse cross-client exploration without compromising privacy. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods on standard benchmarks as well as in sparse and delayed reward settings, achieving both stronger policy personalization and higher sample efficiency.
📝 Abstract
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past experiences while keeping each client's data private during the learning of each client's policy. Many current methods for PFRL rely heavily on exploiting existing reinforcement learning reward signals to derive an optimal policy for each client, thereby neglecting exploration in non-stationary or sparse-reward environments. In this work, we introduce a new exploration-driven framework, Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation (EDPFRL-IM), that leverages an inherent curiosity-driven exploration at each client to promote local exploration and protect client privacy. Furthermore, to facilitate policy discovery via exploration in previously unexplored state spaces, clients add an intrinsic random network distillation (RND) signal to their extrinsic reward. Additionally, the server does not have access to clients' raw experiences or local gradient estimates; instead, the server sends global exploration priors and collects minimal novelty summaries from each client to enable both diverse and coordinated exploration among clients. Experiments in benchmark environments show that our framework outperforms average PFRL benchmarks in policy personalization and sample efficiency, primarily in delayed and sparse reward systems. Overall, EDPFRL-IM enables the integration of a flexible exploratory learning structure into federated reinforcement learning systems while preserving client privacy.
Problem

Research questions and friction points this paper is trying to address.

Personalized Federated Reinforcement Learning
Exploration
Sparse Reward
Non-stationary Environments
Intrinsic Motivation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Personalized Federated Reinforcement Learning
Intrinsic Motivation
Random Network Distillation
Privacy-Preserving Exploration
Exploration-Driven Learning
🔎 Similar Papers
No similar papers found.