Shie Mannor, Yishay Mansour, Aviv Tamar
This appendix covers the ordinary differential equation (ODE) theory necessary for understanding RL convergence proofs. Fundamental ODE definitions and existence/uniqueness results are presented. Linear differential equation systems are analyzed in detail. Asymptotic stability is developed using Lyapunov theory. These results enable the ODE method for analyzing stochastic algorithms. The appendix provides rigorous foundations for convergence proofs of TD learning, Q-learning, and policy gradient algorithms.