科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cambridge University Press eBooks2026-07-31· Reinforcement learning

Reinforcement Learning: Model Based

Shie Mannor, Yishay Mansour, Aviv Tamar

原始摘要(英文原文)· Original abstract
This chapter develops model-based RL approaches where agents explicitly learn MDP transition and reward models from experience. The effective horizon concept quantifies how many steps matter significantly. Off-policy learning with generative models is analyzed first. On-policy learning requires explicit exploration strategies. Several algorithms are presented, including explicit explore-or-exploit, R-MAX and the PAC-MDP framework that ensures probably approximately correct (PAC) policies.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reinforcement Learning: Model Based — 科研速览 Science Skim