Paola Llanos, Javier Russell-Guzmán, Matías Monsalves-Álvarez, Matías Aravena, Martín Maturana, Rodrigo Olivares, Pablo Olivares, José Francisco López-Gil, Rodrigo Yáñez-Sepúlveda, Sonja Buvinic
The minimum Feret diameter is widely used for muscle fiber-size assessment, but identifying discrete size subpopulations remains challenging because conventional thresholds are subjective. We benchmarked CellSize-Clust, an antibody-free clustering framework that integrates eight algorithms with formal K-selection criteria and silhouette voting, using 14,655 fibers from soleus and gastrocnemius muscles of 42 male C57BL/6J mice, segmented with a U-Net deep-learning model from cryosections stained with wheat germ agglutinin (WGA). Internal validation combined Silhouette, Calinski-Harabasz, and Davies-Bouldin indices, while animal-stratified cross-validation and bootstrapping assessed stability. Most criteria supported K = 2, with Agglomerative Clustering with Ward linkage achieving the highest composite score in both muscles and strong agreement with Gaussian Mixture Models (GMM), Spectral Clustering, and K-Means. Cluster proportions were muscle-specific and showed stable cross-validated silhouettes. External validation in 6435 masseter fibers from 12 mice reproduced the two-cluster structure, proportions, and stability, although algorithm rankings did not generalize and the composite score was vulnerable to degenerate partitions. Hartigan's dip test did not reject unimodality, and separation from a unimodal null was limited. CellSize-Clust therefore provides a reproducible method for thresholding a continuous muscle fiber-size distribution rather than evidence of two biologically distinct subpopulations. Associations with myosin heavy chain isoforms require immunohistochemical confirmation. STATEMENT OF SIGNIFICANCE: CellSize-Clust addresses a persistent gap in muscle morphometry: the absence of reproducible, antibody-free frameworks for classifying skeletal muscle fibers by size. Traditional fiber typing relies on costly immunohistochemistry and is subject to inter-laboratory variability. By combining systematic multi-algorithm benchmarking, formal cluster-number justification, animal-level inferential statistics, and fully open data and code, our framework provides a transparent and reproducible pipeline for screening large morphometric datasets and identifying size-based fiber subpopulations before undertaking more resource-intensive immunohistochemical fiber typing. Although CellSize-Clust does not replace myosin heavy chain (MHC) isoform typing - which remains the gold standard - it complements it as a first-pass screening tool, particularly valuable in longitudinal studies of disuse atrophy, obesity, sarcopenia, and neuromuscular disease among others, where thousands of fibers must be phenotyped efficiently.