Institution profile

University of Alabama in Huntsville

Academic institutionnorthamerica · us
Official website
Research library30linked papers
Opportunities0open roles
Selected work

Representative Papers

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

Aug 11, 2026

This work addresses the limitations of existing personalized federated reinforcement learning methods, which overly rely on extrinsic rewards and suffer from insufficient exploration in non-stationary or sparse-reward environments, leading to weak policy personalization and low sample efficiency. To overcome these challenges, we propose the first exploration-driven personalized federated reinforcement learning framework that integrates intrinsic motivation. Specifically, clients leverage Random Network Distillation (RND) to generate intrinsic rewards that enhance local exploration, while the server aggregates only minimal novelty summaries and broadcasts a global exploration prior to coordinate diverse cross-client exploration without compromising privacy. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods on standard benchmarks as well as in sparse and delayed reward settings, achieving both stronger policy personalization and higher sample efficiency.

0 citationsRead paper
Recent publications

Latest Papers

Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation

Aug 11, 2026

This work addresses the limitations of existing personalized federated reinforcement learning methods, which overly rely on extrinsic rewards and suffer from insufficient exploration in non-stationary or sparse-reward environments, leading to weak policy personalization and low sample efficiency. To overcome these challenges, we propose the first exploration-driven personalized federated reinforcement learning framework that integrates intrinsic motivation. Specifically, clients leverage Random Network Distillation (RND) to generate intrinsic rewards that enhance local exploration, while the server aggregates only minimal novelty summaries and broadcasts a global exploration prior to coordinate diverse cross-client exploration without compromising privacy. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods on standard benchmarks as well as in sparse and delayed reward settings, achieving both stronger policy personalization and higher sample efficiency.

0 citationsRead paper