Mohammad Fazle Rabbi
Climate extremes increasingly expose renewable-dominated power systems to simultaneous generation shortfalls and demand surges, creating non-stationary operational conditions that fixed rule-based strategies cannot manage adaptively. This study develops a hybrid physics-informed and data-driven framework in which a Proximal Policy Optimization (PPO) agent coordinates a 500kW proton exchange membrane electrolyser, 1,000 kg compressed hydrogen storage at 350 bar, and a 300kW fuel cell as adaptive grid-resilience infrastructure. The agent processes a 45-dimensional state vector combining meteorological observations, short-term forecasts, LSTM-based climate event indicators, and system states. Systematic benchmarking of five deep reinforcement learning algorithms across 25 hyperparameter configurations identifies PPO with learning rate 3 × 10 − 4 , batch size 2,048, and a 256–128–64 neural architecture as the most stable and sample-efficient controller, converging within 8.9 training hours at sub-3 millisecond inference latency. Multi-regional validation across four European climatic zones over 26,280 simulation hours demonstrates that the learned policy improves resilience by 23.7% and reduces energy not served by 9.6% relative to rule-based control, while lowering average hydrogen inventory by 29.7% and maintaining a 72.5% renewable energy share. Event-stratified analysis reports energy-not-served reductions of 11.0% during cold snaps and 10.0% during heatwaves, with response times accelerated by 45.7% across climate extreme categories. SHAP analysis identifies hydrogen state of charge, renewable generation, and demand as dominant policy drivers; dual ablation experiments quantify the incremental contributions of LSTM event encoding, attention mechanisms, and curriculum learning. A per-node benefit-to-cost ratio of 1.14–1.22 under European Value of Lost Load assumptions confirms economic viability alongside operational performance.