科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cambridge University Press eBooks2026-07-31· Monte Carlo method

Reinforcement Learning: Model Free

Shie Mannor, Yishay Mansour, Aviv Tamar

原始摘要(英文原文)· Original abstract
This chapter presents model-free RL algorithms that learn value functions or policies directly from experience. Q-learning is introduced with updating rules for state–action values. Monte Carlo methods are developed for episodic tasks. The stochastic approximation framework provides mathematical foundations. Temporal difference algorithms bootstrap using value estimates. Q-learning for stochastic MDPs and SARSA for on-policy learning are presented. Multi-step methods interpolate between Monte Carlo and one-step TD.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Reinforcement Learning: Model Free — 科研速览 Science Skim