Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter addresses the curse of dimensionality through value function approximation. Parameterized function approximators represent value functions compactly and generalize across states. The chapter frames approximate policy evaluation as supervised learning. Key theoretical objects include the projection operator and projected Bellman equation. Least squares temporal difference (LSTD) provides a batch solution for linear approximation. Deep Q-networks combining neural networks with experience replay are introduced as the foundation for deep RL successes.