全文下载:
202604015.pdf
文章编号:1672-6987(2026)04-0115-13;DOI:10. 16351/j. 1672-6987. 2026. 04. 015
于哲峰 1 ,王家珂 2 ,王立杰 3* ,刘 洋 2 (1. 鞍山师范大学 物理科学与技术学院,辽宁 鞍山 114007; 2. 青岛科技大学 自动化与电子工程学院,山东 青岛 266061; 3. 青岛大学 自动化学院,山东省工业控制技术重点实验室,山东 青岛 266071)
摘 要:起重机是一种重要的工业设备,已广泛应用于汽车制造、船舶装备等多个领域。针对 具有不确定性的非线性起重机系统,提出一种基于强化学习的优化反步控制方法。首先,引 入辅助变量,将起重机系统重构为严格反馈非线性系统。然后,设计基于神经网络的辨识-执 行网络-评价网络架构,其中辨识、执行网络和评价网络分别用于估计未知动态、实施控制动作 和评估系统性能。进而,在反步法框架下,将虚拟控制器和实际控制输入设计为相应子系统 的优化解,并给出了系统稳定的充分条件。最后,通过仿真验证了算法的有效性。
关键词:起重机系统;优化反步法;强化学习;神经网络
中图分类号:TP 273 文献标志码:A
引用格式:于哲峰,王家珂,王立杰,等 . 基于强化学习的不确定起重机系统定位控制[J]. 青岛科技大学学报(自然科学版),2026,47(4):115-127.
YU Zhefeng, WANG Jiake, WANG Lijie, et al. Reinforcement learning-based uncertain crane system positioning control[J]. Journal of Qingdao University of Science and Technology (Natural Science Edition),2026,47(4):115-127.
Reinforcement Learning-Based Uncertain Crane System Positioning Control
YU Zhefeng1 ,WANG Jiake2 ,WANG Lijie3 ,LIU Yang2 (1. School of Physical Science and Technology, Anshan Normal University, Anshan 114007, China;2. College of Automation and Electronic Engineering, Qingdao University of Science and Technology, Qingdao 266061, China; 3. Shandong Key Laboratory of Industrial Control Technology, School of Automation, Qingdao University, Qingdao 266071, China)
Abstract:Crane is an important industrial equipment and has been widely applied in various fields such as automobile manufacturing and shipbuilding. This paper proposes a reinforcement learning-based optimal backstepping control method for uncertain nonlinear crane systems. Firstly, introduce auxiliary variables to reconstruct the crane system as a strict feedback nonlin⁃ ear system. Then, a neural network-based architecture with identification-execution- network and evaluation-network is designed in this paper. The identification, actor, and critic networks are respectively used to estimate the unknown dynamics, implement control actions, and evaluate system performance. Furthermore, the virtual controller and the actual control input are designed as optimal solutions of the corresponding subsystems, and sufficient conditions for the stability of the system are given in the backstepping framework. Finally, the effectiveness of the algorithm is verified through simulation examples.
Key words:crane system; optimized backstepping; reinforcement learning; neural network
收稿日期:2026-01-10
基金项目:国家自然科学基金项目(62373208,62573249);山东省自然科学基金项目(ZR2024YQ032, ZR2024QF026);山东省泰山学者 项目(tsqn202306218).
作者简介:于哲峰(1973—),男,副教授 . *通信联系人 .