Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter introduces the fundamental Markov decision process model and develops optimal solution methods for finite-horizon problems. Multiple performance criteria are defined, including finite-horizon return and stochastic shortest-path formulations. A key result establishes that Markov policies are sufficient for optimality. The finite-horizon dynamic programming algorithm is derived from the principle of optimality, computing value functions backwards in time via the Bellman equation. The Q-function representation is introduced.