Reza Sepehrzad, Milad Moafi, Nima Khosravi, Ahmed Al durra, Mahdieh S. Sadabadi
The growing dependence on digital communication and machine learning within microgrid (MG) systems has brought new levels of flexibility and automation, but it has also exposed these environments to increasingly coordinated cyber-attacks. This study introduces a defense strategy designed to protect trained control policies during both offline learning and real-time operation by using a Bayesian multi-agent deep reinforcement learning framework. The approach defines a safety region that continuously evaluates policy integrity with respect to load-to-grid (L2G) services and the microgrid's operational goals. In this work, L2G strategies are modeled around energy storage systems (ESS), electric vehicles (EVs), static loads, and renewable energy resources (RER), reflecting the practical structure of modern MGs and their participation in demand response and market-driven energy transactions. To identify harmful data injections and capture the evolving nature of hybrid cyber intrusions, the method combines Gaussian and Bernoulli-based modeling with Euclidean distance measures, enabling accurate detection of abnormal variable behavior. Extensive tests across multiple operating scenarios demonstrate the effectiveness of the framework. The experimental results show an 11.46% reduction in operational costs and prevent roughly $ 6717 in potential attacker gains, underscoring the method's capacity to secure L2G interactions while supporting resilient and economically efficient microgrid operation.