π€ AI Summary
This work addresses the challenge of autonomous control for free-flying robots in microgravity space environments. We propose an end-to-end, six-degree-of-freedom motion control framework based on reinforcement learning (RL), specifically leveraging the Proximal Policy Optimization (PPO) algorithm within an actor-critic architecture. Training is conducted in NVIDIA Isaac Lab under randomized target poses and robot mass parameters. The learned policy is rigorously validated through ground-based testing and, critically, via on-orbit experiments aboard the International Space Station (ISS) using the Astrobee robotβmarking the first deployment of an RL policy for real-space operation. Our approach enables rapid, minute-scale customization of mission-specific behaviors, achieving centimeter-level positional accuracy and sub-degree attitude error. It significantly improves responsiveness and environmental adaptability compared to conventional methods. This work establishes a scalable, simulation-to-reality transferable intelligent control paradigm for future autonomous space operations, on-orbit servicing, and agile orbital logistics.
π Abstract
The US Naval Research Laboratory's (NRL's) Autonomous Planning In-space Assembly Reinforcement-learning free-flYer (APIARY) experiment pioneers the use of reinforcement learning (RL) for control of free-flying robots in the zero-gravity (zero-G) environment of space. On Tuesday, May 27th 2025 the APIARY team conducted the first ever, to our knowledge, RL control of a free-flyer in space using the NASA Astrobee robot on-board the International Space Station (ISS). A robust 6-degrees of freedom (DOF) control policy was trained using an actor-critic Proximal Policy Optimization (PPO) network within the NVIDIA Isaac Lab simulation environment, randomizing over goal poses and mass distributions to enhance robustness. This paper details the simulation testing, ground testing, and flight validation of this experiment. This on-orbit demonstration validates the transformative potential of RL for improving robotic autonomy, enabling rapid development and deployment (in minutes to hours) of tailored behaviors for space exploration, logistics, and real-time mission needs.