科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Cambridge University Press eBooks2026-07-31· Regret

Regret Minimization

Shie Mannor, Yishay Mansour, Aviv Tamar

原始摘要(英文原文)· Original abstract
This chapter studies online learning through the regret minimization framework, focusing on the multi-armed bandit problem. Fundamental lower bounds are established. Improved algorithms are developed including upper confidence bound, achieving optimal regret through optimism in the face of uncertainty. Extensions to MDP settings are discussed. The best-arm identification problem addresses pure exploration with sequential halving, achieving sample-optimal PAC guarantees.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Regret Minimization — 科研速览 Science Skim