Actuator Dynamics Curricula for Narrow-Viability Tasks in Legged Robot Learning
本文针对腿式机器人学习中因早期终止导致梯度信号不足的任务,提出通过逐步降低关节刚度的方法来扩大可行状态集,从而解决此类问题。
本文针对腿式机器人学习中因早期终止导致梯度信号不足的任务,提出通过逐步降低关节刚度的方法来扩大可行状态集,从而解决此类问题。
This paper addresses the challenge of position and attitude control for unmanned aerial systems (UAS) under strong external disturbances. We propose a predictive reinforcement learning (RL) control method integrated with externally triggered signals. Unlike conventional reactive policies, our approach incorporates disturbance-correlated trigger signals as feedforward inputs, enabling the agent to anticipate disturbances and proactively compensate. We innovatively design a proximal policy optimization (PPO)-based deep RL framework that jointly models disturbances, senses trigger events, and conducts high-fidelity simulation training. Experimental results demonstrate that the proposed predictive strategy significantly reduces positional deviation: in simulation, it achieves superior control accuracy and enhanced response proactiveness compared to both baseline controllers and reactive RL methods. To the best of our knowledge, this work is the first to realize trigger-signal-driven, feedforward disturbance compensation within an RL control paradigm.
本文针对腿式机器人学习中因早期终止导致梯度信号不足的任务,提出通过逐步降低关节刚度的方法来扩大可行状态集,从而解决此类问题。
This paper addresses the challenge of position and attitude control for unmanned aerial systems (UAS) under strong external disturbances. We propose a predictive reinforcement learning (RL) control method integrated with externally triggered signals. Unlike conventional reactive policies, our approach incorporates disturbance-correlated trigger signals as feedforward inputs, enabling the agent to anticipate disturbances and proactively compensate. We innovatively design a proximal policy optimization (PPO)-based deep RL framework that jointly models disturbances, senses trigger events, and conducts high-fidelity simulation training. Experimental results demonstrate that the proposed predictive strategy significantly reduces positional deviation: in simulation, it achieves superior control accuracy and enhanced response proactiveness compared to both baseline controllers and reactive RL methods. To the best of our knowledge, this work is the first to realize trigger-signal-driven, feedforward disturbance compensation within an RL control paradigm.