Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models
本文提出一种基于模型的强化学习方法,通过引入近似逆过程模型解决高度灵活模块化制造系统的数据驱动自学习控制问题。
本文提出一种基于模型的强化学习方法,通过引入近似逆过程模型解决高度灵活模块化制造系统的数据驱动自学习控制问题。
In industrial temperature control systems, conventional PID controllers suffer from excessive overshoot, slow settling, and poor adaptability under setpoint step changes and external disturbances. To address these limitations, this paper proposes an event-triggered, dynamic-game-driven real-time self-tuning PID method. The approach introduces a novel event-driven multi-agent game-theoretic learning framework, incorporating automatic boundary detection to accelerate action-space initialization, and provides a rigorous proof of closed-loop system convergence. By synergistically integrating event-triggered control, online reinforcement learning, and dynamic PID gain optimization, the method achieves significant performance improvements: in a printing press temperature loop experiment, overshoot is reduced by 42% and settling time shortened by 37%. Experimental results validate the method’s strong robustness and millisecond-level real-time adaptive capability.
本文提出一种基于模型的强化学习方法,通过引入近似逆过程模型解决高度灵活模块化制造系统的数据驱动自学习控制问题。
In industrial temperature control systems, conventional PID controllers suffer from excessive overshoot, slow settling, and poor adaptability under setpoint step changes and external disturbances. To address these limitations, this paper proposes an event-triggered, dynamic-game-driven real-time self-tuning PID method. The approach introduces a novel event-driven multi-agent game-theoretic learning framework, incorporating automatic boundary detection to accelerate action-space initialization, and provides a rigorous proof of closed-loop system convergence. By synergistically integrating event-triggered control, online reinforcement learning, and dynamic PID gain optimization, the method achieves significant performance improvements: in a printing press temperature loop experiment, overshoot is reduced by 42% and settling time shortened by 37%. Experimental results validate the method’s strong robustness and millisecond-level real-time adaptive capability.