Azam Orooji, Farzaneh Kermani, Seyed Mohsen Hosseini, Alireza Borhani
No clustering algorithm consistently outperformed the others across all medical datasets. Although hierarchical methods showed strong internal and stability performance, these metrics alone were insufficient to predict agreement with clinical labels. Therefore, external validation should complement internal and stability metrics to provide a more comprehensive evaluation of clustering performance in medical data analysis.
BACKGROUND: With the increasing volume and complexity of medical data, clustering methods have become valuable tools for discovering hidden patterns in unlabeled datasets.
OBJECTIVE: This study aims to compare the performance of different clustering algorithms for analyzing medical datasets, and evaluate the effectiveness of different clustering quality criteria, including internal, stability and external metrics.
METHODS: A descriptive analysis was conducted on five publicly available medical datasets: Indian Liver Patient Disease (ILPD), Breast Cancer, Pima Indians Diabetes, Statlog (Heart) and Polycystic Ovary Syndrome (PCOS). After preprocessing (handling missing values, normalization, and variable encoding), eight clustering algorithms-Hierarchical Clustering (HC), k-means clustering, Fuzzy ANalYsis (FANNY), Self-Organizing Tree Algorithm (SOTA), Divisive Analysis (DIANA), Partitioning Around Medoids (PAM), Clustering Large Applications (CLARA) and AGglomerative NESting (AGNES) were applied. Clustering quality was evaluated using internal measures (Connectivity, Silhouette Width, Dunn Index), stability measures (Average Proportion of Non-overlap (APN), Average Distance (AD), Average Distance between Means (ADM) and Figure of Merit (FOM)) and external measures (Accuracy and Adjusted Rand Index (ARI)).
RESULTS: HC and AGNES consistently achieved the best internal and stability performance on the ILPD, Pima, and Statlog datasets, while k-means performed best on the Breast Cancer dataset. On the PCOS dataset, HC, AGNES, k-means, and DIANA showed comparable performance. However, external validation revealed that higher internal and stability scores did not necessarily correspond to greater agreement with clinical class labels, and the optimal algorithm varied across datasets.
CONCLUSION: No clustering algorithm consistently outperformed the others across all medical datasets. Although hierarchical methods showed strong internal and stability performance, these metrics alone were insufficient to predict agreement with clinical labels. Therefore, external validation should complement internal and stability metrics to provide a more comprehensive evaluation of clustering performance in medical data analysis.