科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Biomimetics2026-06-06· Economic dispatch

A Two-Stage PPO–RLMPA Framework for Dynamic Economic Dispatch with Renewable Energy and Storage Integration

Kemal Keskin

原始摘要(英文原文)· Original abstract
The Dynamic Economic Dispatch (DED) problem underpins the cost-efficient and reliable operation of modern power systems, yet valve-point loading, ramp-rate coupling, and the growing share of intermittent wind, photovoltaic, and pumped-storage hydro (PSH) resources render it highly non-convex. Metaheuristic methods typically require large computational budgets and hand-crafted constraint-handling rules, whereas deep reinforcement learning agents rarely guarantee the feasibility of the schedules they produce. To address both limitations, this paper proposes a Two-Stage PPO-RLMPA framework that couples data-driven policy learning with a biomimetic metaheuristic search inspired by marine predator-prey dynamics. In the first stage, a Proximal Policy Optimization (PPO) agent is trained on a Markov Decision Process reformulation of DED in which a deterministic Safety Layer projects every raw action onto the feasible set defined by capacity, ramp-rate, and power-balance constraints, so the policy only observes physically viable transitions. In the second stage, the PPO dispatch is refined by the RLMPA module, a Marine Predators Algorithm (MPA) whose exploration-exploitation balance, Lévy-flight foraging, and Fish Aggregating Devices (FADs) attraction mechanisms emulate strategies documented in marine ecosystems; its step-size factor and FADs probability are further adapted online by a Deep Q-Network. This biomimetics-informed refinement translates predator-prey foraging intelligence into economically efficient thermal dispatch under valve-point non-convexity. Across 30 independent runs on ten- and twenty-unit benchmark systems with wind, PV, and PSH integration, the framework attains best costs of USD 368,763 and USD 737,348 on Test Systems 1 and 2, corresponding to reductions of approximately 1.1% and 4.4% over the CFCEP baseline, with zero post-repair constraint violations in every run.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A Two-Stage PPO–RLMPA Framework for Dynamic Economic Dispatch with Renewable Energy and Storage Integration — 科研速览 Science Skim