Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter introduces reinforcement learning (RL) as the discipline of learning and acting in environments where sequential decisions are made. It traces the origins of RL across multiple disciplines including optimal control, operations research, neuroscience and psychology. The chapter motivates the field through landmark successes in game playing (checkers, backgammon, go, chess), Atari video games, robotics and language model fine-tuning. The primary mathematical model – the Markov Decision Process (MDP) – is introduced as the framework for handling uncertainty in dynamics, actions and knowledge. The chapter outlines the book’s organization into two main themes: planning (optimal decision making with known models) and learning (decision making with unknown models), and presents the key tradeoff between exploitation and exploration inherent to RL problems.