科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cambridge University Press eBooks2026-07-31· Parameterized complexity

Large State Space: Policy Gradient Methods

Shie Mannor, Yishay Mansour, Aviv Tamar

原始摘要(英文原文)· Original abstract
This chapter develops policy gradient methods that directly optimize parameterized policies. The policy performance difference lemma relates performance differences to advantages. The policy gradient theorem enables unbiased gradient estimation from trajectories. REINFORCE implements Monte Carlo gradient estimation. Actor–critic methods reduce variance using learned value function baselines. Proximal policy optimization addresses policy gradient step sizes through trust region methods.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Large State Space: Policy Gradient Methods — 科研速览 Science Skim