Jingying Zhang, Junchao Ren
In this paper, a data-driven iterative Q-learning method was proposed for the safe optimal control problem of constrained nonlinear discrete systems. An augmented utility function was constructed to unify objectives of safety constraints and optimal control by applying a gradient recentered self-concordant barrier function (GRSC-BF) design. It is theoretically proved that the Q-function remains positive definite and converges to the optimum under certain assumptions. An actor-critic framework based on dual back propagation (BP) neural networks was adopted to approximate the iterative control policy and Q-function by constructing a data-driven algorithm without relying on system models. Simulation results show that the proposed method can guarantee the stability and security of control performance in constrained nonlinear systems.