Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter provides a comprehensive treatment of stationary infinite-horizon MDPs with discounted return criterion. The chapter establishes that optimal stationary deterministic policies exist and develops two fundamental algorithms: value iteration and policy iteration. The theoretical foundation rests on contraction mapping theory, with the Banach fixed point theorem guaranteeing unique fixed points and geometric convergence. Error bounds and stopping criteria are derived.