科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Expert Systems with Applications2026-03-03· Computer science

Risk-sensitive actor-critic with static spectral risk measures for online and offline reinforcement learning

Mehrdad Moghimi, Hyejin Ku

原始摘要(英文原文)· Original abstract
• Propose a novel actor-critic framework for static spectral risk measure optimization • Support both online and offline RL with stochastic and deterministic policies • Prove convergence in finite MDPs for the proposed risk-sensitive algorithms • Demonstrate superior performance over baselines in finance, healthcare and robotics • Enable flexible policy tailoring for varying levels of risk aversion The development of Distributional Reinforcement Learning (DRL) has introduced a natural way to incorporate risk sensitivity into value-based and actor-critic methods by employing risk measures other than expectation in the value function. While this approach is widely adopted in many online and offline RL algorithms due to its simplicity, the naive integration of risk measures often results in suboptimal policies. This limitation can be particularly harmful in scenarios where the need for effective risk-sensitive policies is critical and worst-case outcomes carry severe consequences. To address this challenge, we propose a novel framework for optimizing static Spectral Risk Measures (SRM), a flexible family of risk measures that generalizes objectives such as CVaR and Mean-CVaR, and enables the tailoring of risk preferences. Our method is applicable to both online and offline RL algorithms. We establish theoretical guarantees by proving convergence in the finite state-action setting. Moreover, through extensive empirical evaluations, we demonstrate that our algorithms consistently outperform existing risk-sensitive methods in both online and offline environments across diverse domains.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Risk-sensitive actor-critic with static spectral risk measures for online and offline reinforcement learning — 科研速览 Science Skim