Shie Mannor, Yishay Mansour, Aviv Tamar
This chapter develops policy gradient methods that directly optimize parameterized policies. The policy performance difference lemma relates performance differences to advantages. The policy gradient theorem enables unbiased gradient estimation from trajectories. REINFORCE implements Monte Carlo gradient estimation. Actor–critic methods reduce variance using learned value function baselines. Proximal policy optimization addresses policy gradient step sizes through trust region methods.