Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter extends MDP theory to episodic settings where problems continue until reaching goal states. The stochastic shortest path framework is introduced with random termination time and total expected return. The chapter demonstrates that episodic MDPs generalize both finite-horizon and discounted infinite-horizon problems. Modified Bellman equations are derived with value iteration and policy iteration algorithms adapted from the discounted case.