Amreen Batool, Yong-Woon Kim, Yung-Cheol Byun
Wind energy systems are a critical component of global renewable power generation, yet their maintenance remains technically complex and economically demanding, with operation and maintenance (O&M) activities accounting for 20%–35% of the levelized cost of energy. Existing predictive maintenance approaches face three major limitations: reliance on single data modalities, isolated single-agent decision frameworks, and static supervised learning models that struggle to adapt to evolving operational conditions. This paper proposes a Multi-Agent Vision–Language Deep Reinforcement Learning (MAVL-DRL) framework that unifies heterogeneous information sources for coordinated predictive maintenance. The system employs Temporal Convolutional Networks (TCNs) for Supervisory Control and Data Acquisition (SCADA) time-series signals, Distillation with No Labels Version 2 (DINOv2) vision transformers for drone-based blade inspection, Bidirectional Encoder Representations from Transformers (BERT) models for interpreting maintenance logs, and Multi-Layer Perceptrons (MLPs) for meteorological data processing. A cross-attention fusion mechanism learns inter-modal dependencies to construct consistent state representations, which are utilized by a QMIX-based multi-agent reinforcement learning architecture enabling decentralized yet cooperative maintenance policies. Experiments conducted on a real-world 20-turbine offshore wind farm over a 365-day period demonstrate substantial improvements, including 98.3% system availability, only two annual failures (33% fewer than the best-performing baseline), a mean time between failures of 182.5 days, and a 30.8% faster decision response time of 1.8 h. These results position MAVL-DRL as a promising approach for next-generation intelligent predictive maintenance systems in renewable energy infrastructure.