科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Operations Research2026-06-12· Robustness (evolution)

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, Matthieu Geist, Yuejie Chi

原始摘要(英文原文)· Original abstract
Robust Reinforcement Learning Without Compromising Data Efficiency This paper investigates model robustness in reinforcement learning (RL) to reduce the sim-to-real gap in practice. We adopt the framework of distributionally robust Markov decision processes (RMDPs), which aims to learn a policy that optimizes worst-case performance over a prescribed uncertainty set around a nominal Markov decision process (MDP). Despite recent efforts, the sample complexity of RMDPs has remained largely unresolved, leaving open whether distributional robustness has any statistical consequences when benchmarked against standard RL. Assuming access to a generative model of the nominal MDP, the paper provides a near-optimal characterization of the sample complexity of RMDPs across the full range of uncertainty levels under two common uncertainty sets specified by either total variation (TV) distance or chi-squared divergence. Somewhat surprisingly, the results reveal that RMDPs are not necessarily easier or harder to learn than standard MDPs. The statistical consequences of the robustness requirement depend heavily on the size and shape of the uncertainty set, requiring less data in the TV case but more in the chi-squared case compared with standard MDPs.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model — 科研速览 Science Skim