科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Neural networks : the official journal of the International Neural Network Society2026-08-05

APO: Anchored policy optimization by leveraging unsampled actions in continuous spaces.

Weijun Luo, Yingzhuo Liu, Hongsong Tang, Bo Chen, Chaowen Yang, Liuyu Xiang, Zhaofeng He

原始摘要(英文原文)· Original abstract
Policy gradient methods such as Proximal Policy Optimization (PPO) constrain policy updates only on sampled actions, leaving the unsampled action space entirely unconstrained-an issue we term Anchoring Blindness. This limitation induces uncontrolled drift in the policy distribution over unsampled regions, undermining training stability and often leading to suboptimal performance, particularly in continuous action spaces. To address this issue, we propose Anchored Policy Optimization (APO), a PPO variant equipped with Unsampled Action Ratios Regularization (UARR). UARR explicitly constrains the probability ratios between the new and old policies in unsampled action regions, preventing excessive deviation from 1.0. Experimental results on continuous-control tasks demonstrate that APO effectively anchors a broader range of action distributions, significantly improving optimization stability and avoiding convergence to suboptimal solutions. The code is available at https://github.com/wjl-bupt/APO.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

APO: Anchored policy optimization by leveraging unsampled actions in continuous spaces. — 科研速览 Science Skim