科研速览继续刷下去 →
◆ IEEE transactions on cybernetics2026-08-12

Optimal Containment of Multiagent Systems With Multistep Policy Gradient Reinforcement Learning.

Kaitian Chen, Huaicheng Yan, Qiwei Liu, Lingling Lv, Guangjing Song

原始摘要(原文)
This article analyzes the optimal containment control problem of discrete-time multiagent systems (MASs). Multistep temporal difference (TD) learning is integrated with policy gradient (PG) reinforcement learning (RL) to form an online off-policy multistep PG (MS-PG) algorithm. The proposed MS-PG algorithm achieves optimal control performance under completely unknown system dynamics and accommodates asynchronous policy updates among agents. The closed-loop system stability and the algorithmic convergence are rigorously established. Furthermore, an actor-critic neural network (NN) architecture is employed to approximate the control policy and the optimal Q-function, respectively, with data-driven weight update laws derived from the proposed algorithm. To improve training efficiency and sample efficiency, an experience replay (ER) mechanism is incorporated, constructing a hybrid learning framework that fully exploits both offline batch data and online operational data. Finally, simulation results verify the effectiveness of the proposed method.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Optimal Containment of Multiagent Systems With Multistep Policy Gradient Reinforcement Learning. — 科研速览 Science Skim