Shie Mannor, Yishay Mansour, Aviv Tamar
This preface chapter transitions from planning to learning. The chapter distinguishes RL from supervised learning: While supervised learning uses fixed i.i.d. datasets, RL involves an agent interacting online with an environment. Three main learning paradigms are identified: situated agent setting, offline learning and simulation-based learning. The exploration–exploitation tradeoff is emphasized as fundamental. Success metrics including sample complexity bounds, regret minimization and convergence guarantees are discussed.