Jianguo Zhao, Zhijiang Gao, Chunyu Yang, Pak Kin Wong
As one efficient algorithm in adaptive dynamic programming to address adaptive optimal control problems, policy iteration always involves an initial admissible control guess during iteration. Such an initialization process is nontrivial, especially for unstable systems or when system dynamics are totally unknown. To circumvent this limitation, this paper develops a novel adaptive optimal controller design scheme for continuous-time nonlinear systems through neural network-based policy iteration. By employing Carleman linearization, we lift the nonlinear model into a bilinear form, which allows the optimal feedback control to be derived from a state-dependent Riccati (SDR) equation rather than the Hamilton-Jacobi-Bellman equation. To solve the SDR equation under unknown system parameters, we leverage a homotopy-based three-phase policy iteration algorithm that learns the optimal control protocol directly from state and input measurements. In comparison to existing results, the proposed algorithm has the ability to automatically update the initial parameters and specifies the same basis functions of neural networks to approximate all unknown functions, which gives rise to computational improvement. Two case studies are provided to validate our methodology.