Knowledge-Distilled End-to-End Reinforcement Learning for Smooth 6-DOF Thrust Control and Rapid Adaptation to Ocean Currents in Remotely Operated Vehicles

πŸ“… 2026-08-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of achieving low steady-state error, rapid response, high energy efficiency, and smooth thrust simultaneously in end-to-end reinforcement learning for underwater vehicle control under ocean current disturbances. To this end, the authors propose TSRCA-PPO, a method that integrates a multi-objective reward function with a multi-encoder architecture incorporating privileged information. A two-stage knowledge distillation process is employed to train a near-optimal policy capable of producing smooth six-degree-of-freedom thrust commands while rapidly adapting to disturbances. Compared to conventional cascaded P-PID controllers, TSRCA-PPO reduces steady-state position error, attitude error, settling time, energy consumption, and thrust fluctuation by 42.7%, 76.5%, 10.6%, 93.5%, and 15.9%, respectively, thereby significantly enhancing overall control performance.
πŸ“ Abstract
With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicles control. Nevertheless, existing end-to-end reinforcement-learningbased methods still face challenges in achieving optimal control under oceancurrent disturbances. In particular, there remains a lack of a unified control framework that can simultaneously achieve low steady-state tracking error, rapid transient response, energy-efficient operation, and smooth controlforce outputs under disturbances. To address the issue, this paper proposes the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method which learns a near-optimal strategy by a twostage distillation learning framework. The core innovations of this work lie in the reward-function design and the privileged multi-encoder architecture. Ablation studies validate the effectiveness of each module. Simulation results demonstrate that the proposed TSRCA-PPO method consistently outperforms the conventional cascaded P-PID controller across all evaluation metrics. Specifically, TSRCA-PPO reduces the steady-state position error, steady-state attitude error, settling time, energy index, and thrustsmoothness index to 42.7%, 76.5%, 10.6%, 93.5%, and 15.9% of the corresponding P-PID values, respectively.
Problem

Research questions and friction points this paper is trying to address.

ocean current disturbance
6-DOF thrust control
steady-state tracking error
transient response
smooth control
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge distillation
end-to-end reinforcement learning
6-DOF thrust control
ocean current adaptation
smooth control
πŸ’Ό Related Jobs
No related jobs found.
T
Tiankuang Wen
Northwestern Polytechnical University, Xi’an 710129, China
Huiping Li
Huiping Li
School of Marine Science and Technology, Northwestern Polytechnical University
systems and controlmodel predictive controlnetworked control systemsestimation
G
Gang Liu
Wuhan Second Ship Design and Research Institute, Wuhan 430205, China
Y
Yong Jiang
Wuhan Second Ship Design and Research Institute, Wuhan 430205, China