Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter presents model-free RL algorithms that learn value functions or policies directly from experience. Q-learning is introduced with updating rules for state–action values. Monte Carlo methods are developed for episodic tasks. The stochastic approximation framework provides mathematical foundations. Temporal difference algorithms bootstrap using value estimates. Q-learning for stochastic MDPs and SARSA for on-policy learning are presented. Multi-step methods interpolate between Monte Carlo and one-step TD.