科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Biomimetics (Basel, Switzerland)2026-09-07

Online GP-MPC Command Supervision for Robust Reinforcement Learning-Based Quadruped Locomotion.

Seungyeon Lee, Hyunseok Yang

原始摘要(英文原文)· Original abstract
Reinforcement learning-based quadruped locomotion policies can exhibit command-tracking errors under terrain variations and unmodeled dynamics. This study proposes an online bounded Gaussian Process-enhanced model predictive control framework, termed Gaussian Process-Model Predictive Control-Reinforcement Learning(GP-MPC-RL), for command-level supervision of a pretrained locomotion policy. A frozen PPO policy generates the low-level locomotion behavior, while an acados-based MPC supervisor adjusts the velocity command using a nominal command-response model. An online Gaussian Process learns the one-step residual between the nominal prediction and measured robot response, and its uncertainty-weighted forward-velocity correction is incorporated into the MPC prediction. The framework was evaluated in Isaac Lab using a Unitree Go2 quadruped robot model over 20 paired rough-terrain trials at target velocities of 0.3, 0.5, and 0.7 m/s; GP-MPC-RL reduced the mean forward-velocity root mean square error (RMSE) relative to PPO by 38.3%, 21.3%, and 10.1%, respectively. Under a 5 kg payload, GP-MPC-RL reduced velocity RMSE by 29.0% relative to MPC-RL and reduced the 0-5 kg payload-induced degradation by 53.0% (p = 0.019). The average supervisor computation time was 0.112 ms. These results indicate that GP residual adaptation is particularly effective when the nominal command-response model becomes inaccurate, improving robustness without retraining the underlying locomotion policy.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Online GP-MPC Command Supervision for Robust Reinforcement Learning-Based Quadruped Locomotion. — 科研速览 Science Skim