Jiajie Li, Jia Xu, Yongjun Lian, Honglu Zhu
The effective deployment of intelligent algorithms, including PV station efficiency assessment and fault diagnosis model training, relies on high-quality operational datasets with reliable labeling. Therefore, the efficient and systematic partitioning of massive PV array operational data is of significant engineering application value. To achieve effective partitioning of PV array anomaly datasets under different states, this paper proposes a hybrid method integrating unsupervised clustering and statistical analysis. First, construct the output time-series feature indices of the PV array. Then utilize the K-means algorithm to perform preliminary clustering on these features to identify and separate short-circuit and open-circuit datasets. Further apply Fuzzy C-Means clustering to the statistical features of the PV array operational data to identify PV strings under partial shading conditions, thereby partitioning the shading dataset. This proposed method, based on time-series and statistical distribution features, can effectively categorize datasets into normal, open-circuit, short-circuit, and partial shading classes. Experimental results demonstrate that the proposed two-stage anomaly dataset partitioning method achieves an accuracy of 99.8%. This represents an improvement of over 30% compared to traditional single-clustering methods based solely on time-series features, providing a reliable and effective new approach for PV array anomaly dataset partitioning.