Jianmin Wu, Hao Shen
The Kleinman iteration is a policy iteration method for solving Riccati equations and forms the basis of many reinforcement learning (RL) algorithms. However, its direct application to inverse RL problems is limited by its reliance on known cost functions. To overcome this limitation, we propose a modified Kleinman iteration algorithm that simultaneously estimates the cost weight and the value function matrix during policy evaluation without an additional update step for the cost weight, thereby improving computational efficiency and reducing the number of iterations. Moreover, we develop a data-driven inverse RL algorithm by introducing a piecewise constant, persistently exciting input during the data collection phase. This design avoids the use of Kronecker product operations by reducing the dimensionality of the matrices. The convergence properties and solution non-uniqueness of the proposed method are rigorously analyzed. Finally, the effectiveness of the proposed algorithms is demonstrated through simulations on a power system model, with performance comparisons against existing approaches.