Enhancing efficiency and propulsion in bio-mimetic robotic fish through end-to-end deep reinforcement learning

📅 2024-03-01
🏛️ The Physics of Fluids
📈 Citations: 9
Influential: 0
📄 PDF
🤖 AI Summary
Bionic robotic fish suffer from low propulsion efficiency and high energy consumption. Method: This study proposes an end-to-end deep reinforcement learning (DRL) control framework, introducing— for the first time in underwater bionic robotics—extended pressure sensing combined with temporal Transformer modeling, integrated with a policy transfer mechanism to enhance training stability and environmental adaptability. Training achieves autonomous, stable, and rapid convergence within CFD simulations (Re = 6000). Contribution/Results: The DRL policy improves propulsion efficiency by 37% and reduces specific energy consumption per unit thrust by 29% over conventional pre-programmed gaits. Flow-field analysis reveals that efficiency stems from embodied regulation of body deformation and vortex–body interactions. The core contribution is a novel bio-inspired locomotion control paradigm unifying perception, spatiotemporal modeling, and decision-making.

Technology Category

Application Category

📝 Abstract
Aquatic organisms are known for their ability to generate efficient propulsion with low energy expenditure. While existing research has sought to leverage bio-inspired structures to reduce energy costs in underwater robotics, the crucial role of control policies in enhancing efficiency has often been overlooked. In this study, we optimize the motion of a bio-mimetic robotic fish using deep reinforcement learning (DRL) to maximize propulsion efficiency and minimize energy consumption. Our novel DRL approach incorporates extended pressure perception, a transformer model processing sequences of observations, and a policy transfer scheme. Notably, significantly improved training stability and speed within our approach allow for end-to-end training of the robotic fish. This enables agiler responses to hydrodynamic environments and possesses greater optimization potential compared to pre-defined motion pattern controls. Our experiments are conducted on a serially connected rigid robotic fish in a free stream with a Reynolds number of 6000 using computational fluid dynamics simulations. The DRL-trained policies yield impressive results, demonstrating both high efficiency and propulsion. The policies also showcase the agent's embodiment, skillfully utilizing its body structure and engaging with surrounding fluid dynamics, as revealed through flow analysis. This study provides valuable insights into the bio-mimetic underwater robots optimization through DRL training, capitalizing on their structural advantages, and ultimately contributing to more efficient underwater propulsion systems.
Problem

Research questions and friction points this paper is trying to address.

Optimizing bio-mimetic robotic fish motion for efficiency
Enhancing propulsion with deep reinforcement learning
Minimizing energy consumption in underwater robotics
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-end deep reinforcement learning optimization
Transformer model for observation sequence processing
Policy transfer scheme for enhanced training stability
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xinyu Cui
Institute of Automation, Chinese Academy of Science, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
B
Boai Sun
Zhejiang University-Westlake University Joint Training, Zhejiang University, Hangzhou 310027, China
Y
Yi Zhu
Key Laboratory of Coastal Environment and Resources of Zhejiang Province, School of Engineering, Westlake University, Hangzhou 310030, China; Institute of Advanced Technology, Westlake Institute for Advanced Study, Hangzhou 310024, China
N
Ning Yang
Institute of Automation, Chinese Academy of Science, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
H
Haifeng Zhang
Institute of Automation, Chinese Academy of Science, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
W
Weicheng Cui
Key Laboratory of Coastal Environment and Resources of Zhejiang Province, School of Engineering, Westlake University, Hangzhou 310030, China; Institute of Advanced Technology, Westlake Institute for Advanced Study, Hangzhou 310024, China
D
Dixia Fan
Key Laboratory of Coastal Environment and Resources of Zhejiang Province, School of Engineering, Westlake University, Hangzhou 310030, China; Institute of Advanced Technology, Westlake Institute for Advanced Study, Hangzhou 310024, China
J
Jun Wang
Department of Computer Science, University College London, London WC1E 6BT, United Kingdom