Jinli Yao, Yong Zeng
Clustering is a fundamental technique in unsupervised learning, enabling the discovery of patterns and natural groupings in data without prior labels. Despite its widespread applications across domains, the field of clustering faces persistent challenges, including a lack of universally accepted definitions, inconsistent classification criteria, and varying evaluation metrics. This review paper addresses these gaps by exploring the core question: What defines a good cluster? We investigate and summarize the induction principle behind clustering problems, clustering algorithms, and evaluation indices. The paper classifies clustering algorithms based on their criteria and principles, providing a structured understanding of their methodologies. It further categorizes datasets into synthetic and real-world examples, identifying the challenges posed by diverse cluster characteristics, such as varying shapes, densities, sizes, and overlapping cases, alongside high-dimensionality. A comprehensive review of evaluation indices-grouped into compactness, connectedness, and separation types-highlights their importance in assessing clustering quality. By consolidating these aspects, this review provides a cohesive framework to understand clustering principles and their applications.