Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter develops model-based RL approaches where agents explicitly learn MDP transition and reward models from experience. The effective horizon concept quantifies how many steps matter significantly. Off-policy learning with generative models is analyzed first. On-policy learning requires explicit exploration strategies. Several algorithms are presented, including explicit explore-or-exploit, R-MAX and the PAC-MDP framework that ensures probably approximately correct (PAC) policies.